AmberYifan/capsdnum-marin-8b-base-math_random_b8000_s0
The AmberYifan/capsdnum-marin-8b-base-math_random_b8000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. It was trained on the capsd_marin-8b-base-n80000-numina__mix_math_random_b8000_s0 dataset, suggesting a specialization in mathematical or numerical reasoning tasks. With a context length of 8192 tokens, this model is designed for applications requiring processing and generating content related to its specific fine-tuning domain.
Loading preview...
Model Overview
This model, marin-8b-base_math_random_b8000_s0, is an 8 billion parameter language model developed by AmberYifan. It is a fine-tuned variant of the marin-community/marin-8b-base architecture, specifically adapted through training on the capsd_marin-8b-base-n80000-numina__mix_math_random_b8000_s0 dataset. The fine-tuning process involved a learning rate of 1e-05, a total training batch size of 64, and utilized a cosine learning rate scheduler over 1 epoch.
Key Characteristics
- Base Model: Fine-tuned from
marin-community/marin-8b-base. - Parameter Count: 8 billion parameters.
- Context Length: Supports an 8192-token context window.
- Training Data: Specialized training on the
capsd_marin-8b-base-n80000-numina__mix_math_random_b8000_s0dataset, indicating a focus on mathematical or numerical reasoning.
Training Details
The model was trained using the following key hyperparameters:
- Learning Rate: 1e-05
- Optimizer: ADAMW_TORCH
- Batch Size: Total train batch size of 64 (2 per device with 8 gradient accumulation steps on 4 GPUs).
- Epochs: 1
- Scheduler: Cosine LR scheduler with 0.03 warmup steps.
Potential Use Cases
Given its fine-tuning dataset, this model is likely suitable for tasks that involve:
- Mathematical problem-solving.
- Numerical analysis and generation.
- Applications requiring reasoning over structured numerical data.