AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b4000_s0
AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b4000_s0 is a 1.7 billion parameter language model based on the Qwen3 architecture, fine-tuned from Qwen/Qwen3-1.7B-Base. This model is specifically optimized for mathematical tasks, having been trained on a specialized dataset for improved performance in this domain. With a context length of 32768 tokens, it is designed for applications requiring robust mathematical reasoning and processing.
Loading preview...
Model Overview
This model, AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b4000_s0, is a specialized variant of the Qwen3-1.7B-Base architecture. It has been fine-tuned from the original Qwen/Qwen3-1.7B-Base model, focusing on enhancing its capabilities in mathematical reasoning and problem-solving.
Key Characteristics
- Base Model: Qwen3-1.7B-Base
- Parameter Count: 1.7 billion parameters
- Context Length: 32768 tokens
- Specialization: Optimized for mathematical tasks through fine-tuning on the
capsd_Qwen3-1.7B-Base-n10000__mix_math_ppl_b4000_s0dataset.
Training Details
The model underwent a single epoch of training with a learning rate of 1e-05 and a total batch size of 64. It utilized a cosine learning rate scheduler with 0.03 warmup steps. The training was conducted using Transformers 5.7.0 and Pytorch 2.13.0+cu130.
Good For
- Applications requiring a compact yet capable model for mathematical computations.
- Tasks involving numerical reasoning, equation solving, or data analysis where mathematical proficiency is crucial.
- Scenarios where the Qwen3 architecture's base capabilities are desired, with an added emphasis on mathematical performance.