AmberYifan/capsd-marin-8b-base-math_kcenter_b2000_s0
The AmberYifan/capsd-marin-8b-base-math_kcenter_b2000_s0 model is an 8 billion parameter language model, fine-tuned from marin-community/marin-8b-base. This model is specifically optimized for mathematical tasks, having been trained on the capsd_marin-8b-base-n80000-numina__mix_math_kcenter_b2000_s0 dataset. It features an 8192 token context length, making it suitable for applications requiring robust mathematical reasoning and problem-solving capabilities.
Loading preview...
Model Overview
The AmberYifan/capsd-marin-8b-base-math_kcenter_b2000_s0 is an 8 billion parameter language model, fine-tuned from the marin-community/marin-8b-base architecture. This model has been specifically adapted for enhanced performance in mathematical domains.
Key Capabilities
- Mathematical Reasoning: Optimized through fine-tuning on a specialized dataset, making it suitable for tasks involving mathematical problem-solving and understanding.
- Base Model: Built upon the
marin-8b-basemodel, suggesting a strong foundation for general language understanding prior to its mathematical specialization. - Context Length: Supports an 8192-token context window, allowing for processing and generating longer sequences of text, which can be beneficial for complex mathematical problems requiring detailed context.
Training Details
The model was fine-tuned using the capsd_marin-8b-base-n80000-numina__mix_math_kcenter_b2000_s0 dataset. Key training hyperparameters included a learning rate of 1e-05, a total batch size of 64 (with gradient accumulation), and a cosine learning rate scheduler with 0.03 warmup steps over 1 epoch. The training utilized Transformers 5.7.0 and Pytorch 2.13.0+cu130.
When to Use This Model
This model is particularly well-suited for applications that require:
- Mathematical Problem Solving: Ideal for tasks where accurate mathematical computation and reasoning are critical.
- Specialized Domain Applications: Use cases that benefit from a model specifically trained on mathematical datasets, potentially outperforming general-purpose models in this area.