AmberYifan/capsd-marin-8b-base-math_ifd_b4000_s0
The AmberYifan/capsd-marin-8b-base-math_ifd_b4000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model is specifically adapted for mathematical instruction-following tasks, leveraging the capsd_marin-8b-base-n80000-numina__mix_math_ifd_b4000_s0 dataset. It is designed to enhance performance in mathematical reasoning and problem-solving contexts.
Loading preview...
Model Overview
This model, AmberYifan/capsd-marin-8b-base-math_ifd_b4000_s0, is an 8 billion parameter language model. It is a fine-tuned variant of the marin-community/marin-8b-base architecture, specifically optimized for mathematical tasks.
Key Characteristics
- Base Model: Fine-tuned from
marin-community/marin-8b-base. - Parameter Count: 8 billion parameters.
- Context Length: Supports an 8192-token context window.
- Specialization: The model has undergone fine-tuning on the
capsd_marin-8b-base-n80000-numina__mix_math_ifd_b4000_s0dataset, indicating a focus on mathematical instruction-following and problem-solving.
Training Details
The fine-tuning process utilized the following hyperparameters:
- Learning Rate: 1e-05
- Batch Sizes:
train_batch_sizeof 2,eval_batch_sizeof 8, with agradient_accumulation_stepsof 8, resulting in atotal_train_batch_sizeof 64. - Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08.
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps.
- Epochs: Trained for 1 epoch.
Potential Use Cases
Given its fine-tuning on a mathematical dataset, this model is likely suitable for applications requiring:
- Mathematical problem-solving.
- Instruction following in quantitative domains.
- Generating or understanding mathematical explanations.