AmberYifan/capsdnum-marin-8b-base-math_cap_b8000_s0
AmberYifan/capsdnum-marin-8b-base-math_cap_b8000_s0 is an 8 billion parameter language model fine-tuned from marin-community/marin-8b-base. This model is specifically trained on the capsd_marin-8b-base-n80000-numina__mix_math_cap_b8000_s0 dataset, suggesting an optimization for mathematical reasoning or numerical tasks. It features a context length of 8192 tokens, making it suitable for processing moderately long inputs in its specialized domain.
Loading preview...
Model Overview
AmberYifan/capsdnum-marin-8b-base-math_cap_b8000_s0 is an 8 billion parameter language model derived from the marin-community/marin-8b-base architecture. This model has undergone a specific fine-tuning process using the capsd_marin-8b-base-n80000-numina__mix_math_cap_b8000_s0 dataset.
Training Details
The fine-tuning process utilized the following key hyperparameters:
- Learning Rate: 1e-05
- Batch Sizes: A
train_batch_sizeof 2 andeval_batch_sizeof 8, with atotal_train_batch_sizeof 64 andtotal_eval_batch_sizeof 32, achieved through agradient_accumulation_stepsof 8. - Optimizer: ADAMW_TORCH with default betas and epsilon.
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps.
- Epochs: Trained for 1 epoch.
The training was conducted on a multi-GPU setup with 4 devices, using Transformers 5.7.0, Pytorch 2.13.0+cu130, Datasets 4.0.0, and Tokenizers 0.22.2.
Potential Use Cases
Given its specific training dataset, this model is likely optimized for tasks involving:
- Mathematical problem-solving
- Numerical reasoning
- Processing and generating content related to quantitative data