AmberYifan/capsd-Qwen3-1.7B-Base-math_cap_b4000_s0
AmberYifan/capsd-Qwen3-1.7B-Base-math_cap_b4000_s0 is a 2 billion parameter Qwen3-1.7B-Base model fine-tuned by AmberYifan. This model is specifically optimized for mathematical tasks, having been trained on the capsd_Qwen3-1.7B-Base-n10000__mix_math_cap_b4000_s0 dataset. It is designed to enhance performance in mathematical reasoning and problem-solving within its 32768 token context length.
Loading preview...
Model Overview
AmberYifan/capsd-Qwen3-1.7B-Base-math_cap_b4000_s0 is a specialized 2 billion parameter language model, fine-tuned from the Qwen/Qwen3-1.7B-Base architecture. Its primary differentiation lies in its targeted training on a mathematical dataset, specifically capsd_Qwen3-1.7B-Base-n10000__mix_math_cap_b4000_s0.
Key Capabilities
- Mathematical Task Optimization: The model has undergone specific fine-tuning to improve its performance on mathematical reasoning and problem-solving tasks.
- Base Model Architecture: Built upon the Qwen3-1.7B-Base, it inherits the foundational capabilities of the Qwen3 series.
- Context Length: Supports a substantial context window of 32768 tokens, beneficial for complex problems requiring extensive context.
Training Details
The model was trained with a learning rate of 1e-05, using an AdamW optimizer and a cosine learning rate scheduler. Training involved a total batch size of 64 across 4 devices for 1 epoch. The training environment utilized Transformers 5.7.0 and Pytorch 2.13.0+cu130.
Good For
- Applications requiring enhanced mathematical understanding and generation.
- Research and development in mathematical AI.
- Scenarios where a compact yet mathematically capable model is needed.