AmberYifan/capsdnum-marin-8b-base-math_cap_b1000_s0
The AmberYifan/capsdnum-marin-8b-base-math_cap_b1000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. It was specifically trained on the capsd_marin-8b-base-n80000-numina__mix_math_cap_b1000_s0 dataset, suggesting an optimization for mathematical reasoning tasks. This model is intended for applications requiring enhanced numerical and mathematical capabilities within an 8192-token context window.
Loading preview...
Model Overview
The AmberYifan/capsdnum-marin-8b-base-math_cap_b1000_s0 is an 8 billion parameter language model, derived from the marin-community/marin-8b-base architecture. It has been fine-tuned with a specific focus on mathematical capabilities, utilizing the capsd_marin-8b-base-n80000-numina__mix_math_cap_b1000_s0 dataset.
Training Details
The model underwent a fine-tuning process with the following key hyperparameters:
- Learning Rate: 1e-05
- Batch Sizes: A
train_batch_sizeof 2 andeval_batch_sizeof 8, with atotal_train_batch_sizeof 64 andtotal_eval_batch_sizeof 32, achieved through gradient accumulation over 8 steps. - Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08.
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps.
- Epochs: Trained for 1 epoch.
Intended Use Cases
Given its specialized training on a mathematical dataset, this model is likely best suited for:
- Mathematical Reasoning: Tasks involving numerical problem-solving, equation handling, and quantitative analysis.
- Scientific Applications: Potentially useful in fields requiring precise numerical understanding and generation.
Limitations
As indicated in the original model card, further information regarding specific limitations and broader intended uses is needed for a comprehensive understanding of its scope and performance boundaries.