AmberYifan/capsdnum-marin-8b-base-math_random_b4000_s0
AmberYifan/capsdnum-marin-8b-base-math_random_b4000_s0 is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model was specifically trained on a dataset focused on mathematical and random number generation tasks. It is intended for applications requiring specialized performance in numerical reasoning and related computational problems.
Loading preview...
Model Overview
This model, AmberYifan/capsdnum-marin-8b-base-math_random_b4000_s0, is an 8 billion parameter language model derived from the marin-community/marin-8b-base architecture. It has been fine-tuned with a specific focus on mathematical and random number generation tasks, utilizing the capsd_marin-8b-base-n80000-numina__mix_math_random_b4000_s0 dataset.
Key Training Details
The fine-tuning process involved a single epoch with a learning rate of 1e-05. Training was conducted on a multi-GPU setup (4 devices) with a total batch size of 64, achieved through a train_batch_size of 2 and gradient_accumulation_steps of 8. The optimizer used was ADAMW_TORCH with standard betas and epsilon, and a cosine learning rate scheduler with 0.03 warmup steps.
Potential Use Cases
Given its specialized training, this model is likely best suited for:
- Mathematical problem-solving: Tasks involving numerical reasoning, calculations, or mathematical concept understanding.
- Random number generation contexts: Applications where the model needs to process or generate sequences related to randomness.
Further information regarding specific capabilities, limitations, and detailed evaluation data is not provided in the current model card.