AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b4000_s0
AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b4000_s0 is a 4 billion parameter language model fine-tuned from Qwen/Qwen3-4B-Base. This model has been specifically adapted using the capsd_Qwen3-4B-Base-n80000-numina__mix_math_ppl_b4000_s0 dataset, indicating a specialization in mathematical reasoning or related tasks. With a context length of 32768 tokens, it is designed for applications requiring processing of extensive numerical or technical information.
Loading preview...
Model Overview
This model, AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_ppl_b4000_s0, is a specialized version of the 4 billion parameter Qwen3-4B-Base model. It has undergone fine-tuning on a specific dataset, capsd_Qwen3-4B-Base-n80000-numina__mix_math_ppl_b4000_s0, suggesting an optimization for tasks involving mathematical processing or numerical understanding.
Training Details
The fine-tuning process utilized a learning rate of 1e-05, a total batch size of 64 (with a train batch size of 2 and gradient accumulation steps of 8), and was trained for 1 epoch. The optimizer used was AdamW with cosine learning rate scheduling and a warmup of 0.03. The training was conducted across 4 GPUs.
Key Characteristics
- Base Model: Qwen3-4B-Base
- Parameter Count: 4 billion
- Context Length: 32768 tokens
- Specialization: Fine-tuned on a dataset (
capsd_Qwen3-4B-Base-n80000-numina__mix_math_ppl_b4000_s0) that implies a focus on mathematical or numerical tasks.
Potential Use Cases
Given its fine-tuning on a math-related dataset, this model is likely suitable for applications requiring:
- Mathematical problem-solving
- Numerical data analysis
- Tasks involving quantitative reasoning