AmberYifan/capsd-marin-8b-base-math_less_b4000_s0
The AmberYifan/capsd-marin-8b-base-math_less_b4000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model was trained on the capsd_marin-8b-base-n80000-numina__mix_math_less_b4000_s0 dataset, suggesting a specialization in mathematical reasoning or related tasks. It leverages a cosine learning rate scheduler and AdamW optimizer, indicating a focus on robust and efficient training for its specific domain.
Loading preview...
Model Overview
The AmberYifan/capsd-marin-8b-base-math_less_b4000_s0 is an 8 billion parameter language model, fine-tuned from the marin-community/marin-8b-base architecture. This model was specifically trained on the capsd_marin-8b-base-n80000-numina__mix_math_less_b4000_s0 dataset, which implies a targeted optimization for tasks involving mathematical reasoning or quantitative analysis.
Training Details
The fine-tuning process utilized the following key hyperparameters:
- Learning Rate: 1e-05
- Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08
- Batch Size: A total training batch size of 64 (2 per device across 4 GPUs with 8 gradient accumulation steps)
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps
- Epochs: Trained for 1 epoch
This configuration suggests a focused and efficient fine-tuning approach to adapt the base model to its specialized dataset.
Intended Use Cases
While specific intended uses and limitations are not detailed in the provided README, the dataset name strongly indicates that this model is likely optimized for:
- Mathematical problem-solving
- Quantitative reasoning tasks
- Applications requiring numerical understanding