AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b1000_s0
AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b1000_s0 is a 4 billion parameter language model based on the Qwen3-4B-Base architecture, fine-tuned for mathematical tasks. This model is specifically optimized using the capsd_Qwen3-4B-Base-n80000-numina__mix_math_ppl_b1000_s0 dataset, suggesting enhanced performance in mathematical problem-solving and reasoning. It leverages a 32K context length, making it suitable for processing longer mathematical queries and related text.
Loading preview...
Model Overview
This model, AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b1000_s0, is a specialized version of the Qwen3-4B-Base architecture, featuring 4 billion parameters and a 32K context length. It has been fine-tuned on a specific dataset, capsd_Qwen3-4B-Base-n80000-numina__mix_math_ppl_b1000_s0, indicating a focus on improving its capabilities in mathematical domains.
Key Characteristics
- Base Model: Qwen3-4B-Base
- Parameter Count: 4 billion
- Context Length: 32,768 tokens
- Fine-tuning Focus: Enhanced performance on mathematical tasks, as suggested by the training dataset.
Training Details
The model underwent a single epoch of training with a learning rate of 1e-05, using an AdamW optimizer and a cosine learning rate scheduler. Training involved a total batch size of 64 across 4 GPUs, with a gradient accumulation of 8 steps. This configuration aims to refine the model's understanding and generation for math-related content.
Intended Use Cases
This model is likely best suited for applications requiring strong mathematical reasoning or processing of numerical and scientific text. Its fine-tuning on a math-specific dataset suggests improved accuracy and relevance for such tasks compared to general-purpose models.