AmberYifan/capsd-Qwen3-1.7B-Base-math_cap_b2000_s0
AmberYifan/capsd-Qwen3-1.7B-Base-math_cap_b2000_s0 is a 1.7 billion parameter Qwen3-based language model fine-tuned from Qwen/Qwen3-1.7B-Base. This model is specifically fine-tuned on the capsd_Qwen3-1.7B-Base-n10000__mix_math_cap_b2000_s0 dataset, indicating an optimization for mathematical capabilities. It is designed for tasks requiring numerical reasoning and problem-solving, leveraging a 32768-token context length.
Loading preview...
Model Overview
AmberYifan/capsd-Qwen3-1.7B-Base-math_cap_b2000_s0 is a 1.7 billion parameter language model based on the Qwen3 architecture. It is a fine-tuned variant of the original Qwen/Qwen3-1.7B-Base model.
Key Characteristics
- Base Model: Qwen3-1.7B-Base
- Parameter Count: 1.7 billion parameters
- Context Length: 32768 tokens
- Fine-tuning Dataset:
capsd_Qwen3-1.7B-Base-n10000__mix_math_cap_b2000_s0
Training Details
The model was trained with the following hyperparameters:
- Learning Rate: 1e-05
- Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08
- Scheduler: Cosine learning rate scheduler with 0.03 warmup steps
- Epochs: 1
- Batch Size: A total training batch size of 64 (2 per device across 4 GPUs with 8 gradient accumulation steps).
Intended Use
While specific intended uses and limitations require more information, the fine-tuning dataset suggests a focus on enhancing the model's capabilities in mathematical reasoning and problem-solving tasks. Developers should consider this model for applications where numerical accuracy and mathematical understanding are critical.