AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b8000_s0
AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b8000_s0 is a 4 billion parameter Qwen3-Base model, fine-tuned by AmberYifan, specifically optimized for mathematical capabilities. This model leverages a 32K context length and is trained on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b8000_s0 dataset. Its primary strength lies in enhanced performance for mathematical reasoning and problem-solving tasks.
Loading preview...
Model Overview
AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b8000_s0 is a 4 billion parameter language model, fine-tuned from the base Qwen/Qwen3-4B-Base architecture. This model has been specifically adapted using the capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b8000_s0 dataset, indicating a focus on improving its mathematical reasoning and problem-solving abilities.
Training Details
The model underwent a single epoch of fine-tuning with a learning rate of 1e-05, utilizing an AdamW optimizer. Training was conducted across 4 devices with a total batch size of 64, employing a cosine learning rate scheduler with 0.03 warmup steps. The training environment included Transformers 5.8.0, Pytorch 2.13.0+cu130, Datasets 4.0.0, and Tokenizers 0.22.2.
Intended Use Cases
This model is particularly suited for applications requiring enhanced mathematical understanding and computation. Its fine-tuning on a specialized mathematical dataset suggests improved performance in tasks such as:
- Solving mathematical problems.
- Generating mathematical explanations.
- Assisting with quantitative analysis.