AmberYifan/capsd-marin-8b-base-math_dsir_b4000_s0
The AmberYifan/capsd-marin-8b-base-math_dsir_b4000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. It was specifically trained on the capsd_marin-8b-base-n80000-numina__mix_math_dsir_b4000_s0 dataset, suggesting an optimization for mathematical reasoning tasks. This model is intended for applications requiring specialized performance in mathematical domains.
Loading preview...
Model Overview
The AmberYifan/capsd-marin-8b-base-math_dsir_b4000_s0 is an 8 billion parameter language model derived from the marin-community/marin-8b-base architecture. This model has undergone a specific fine-tuning process using the capsd_marin-8b-base-n80000-numina__mix_math_dsir_b4000_s0 dataset.
Training Details
The fine-tuning procedure involved the following hyperparameters:
- Learning Rate: 1e-05
- Batch Size: A
train_batch_sizeof 2 andeval_batch_sizeof 8, leading to atotal_train_batch_sizeof 64 andtotal_eval_batch_sizeof 32, utilizing 4 GPUs. - Optimizer: ADAMW_TORCH with default betas and epsilon.
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps.
- Epochs: Trained for 1 epoch.
Intended Use
While specific intended uses and limitations are not detailed in the provided information, the dataset name strongly implies that this model is specialized for tasks involving mathematical reasoning and problem-solving. Developers should consider its fine-tuning on a math-centric dataset when evaluating its suitability for their applications.