AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b2000_s0
The AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b2000_s0 model is a 1.7 billion parameter language model, fine-tuned by AmberYifan from the Qwen3-1.7B-Base architecture. This model is specifically optimized for mathematical tasks, leveraging a specialized dataset for improved performance in this domain. With a context length of 32768 tokens, it is designed for applications requiring robust mathematical reasoning and processing.
Loading preview...
Model Overview
AmberYifan/capsd-Qwen3-1.7B-Base-math_ppl_b2000_s0 is a 1.7 billion parameter language model, fine-tuned from the base Qwen/Qwen3-1.7B-Base architecture. This model has been specialized through fine-tuning on the capsd_Qwen3-1.7B-Base-n10000__mix_math_ppl_b2000_s0 dataset, indicating a strong focus on mathematical capabilities.
Key Characteristics
- Base Model: Qwen3-1.7B-Base
- Parameter Count: 1.7 billion
- Context Length: 32768 tokens
- Specialization: Fine-tuned for mathematical tasks, suggesting enhanced performance in numerical reasoning and problem-solving.
Training Details
The model underwent a single epoch of training with a learning rate of 1e-05, using a total batch size of 64 across 4 GPUs. The optimizer used was AdamW with cosine learning rate scheduling. This focused training approach on a specialized dataset aims to improve its proficiency in mathematical domains.
Intended Use Cases
This model is likely suitable for applications requiring:
- Mathematical problem-solving
- Numerical analysis
- Tasks involving quantitative reasoning
Due to its specific fine-tuning, it is expected to perform well in scenarios where mathematical accuracy and understanding are critical.