AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b2000_s0
AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b2000_s0 is a 4 billion parameter language model, fine-tuned from Qwen/Qwen3-4B-Base, with a 32768 token context length. This model has been specifically fine-tuned on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b2000_s0 dataset. It is optimized for tasks related to mathematical capabilities, leveraging its base Qwen3 architecture.
Loading preview...
Overview
This model, capsd-qwen3-numina-Qwen3-4B-Base-math_cap_b2000_s0, is a 4 billion parameter language model derived from the Qwen3-4B-Base architecture. It features a substantial context length of 32768 tokens, making it suitable for processing longer inputs.
Key Capabilities
- Mathematical Fine-tuning: The model has undergone specific fine-tuning on the
capsd_Qwen3-4B-Base-n80000-numina__mix_math_cap_b2000_s0dataset, indicating an optimization for tasks requiring mathematical understanding or generation. - Base Model: Built upon the robust Qwen3-4B-Base, it inherits the foundational capabilities of the Qwen3 series.
Training Details
The fine-tuning process involved a learning rate of 1e-05, a total training batch size of 64, and utilized 4 devices with a gradient accumulation of 8 steps. The training was conducted for 1 epoch using an AdamW optimizer and a cosine learning rate scheduler with 0.03 warmup steps. The training environment included Transformers 5.8.0, Pytorch 2.13.0+cu130, Datasets 4.0.0, and Tokenizers 0.22.2.
Intended Use Cases
This model is primarily intended for applications and research focusing on enhancing mathematical reasoning and problem-solving capabilities within the Qwen3-4B-Base framework.