Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_rel_1e-1_alt_1_per_5_1p0_0p0_1p0_grpo_42_rule
Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_rel_1e-1_alt_1_per_5_1p0_0p0_1p0_grpo_42_rule is a 1.7 billion parameter language model fine-tuned from Qwen/Qwen3-1.7B-Base. Developed by Kazuki1450, this model was trained using the TRL framework and incorporates the GRPO method. It is specifically optimized for mathematical reasoning tasks, leveraging techniques from the DeepSeekMath research. This model is suitable for applications requiring enhanced mathematical problem-solving capabilities.
Loading preview...
Model Overview
This model, Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_rel_1e-1_alt_1_per_5_1p0_0p0_1p0_grpo_42_rule, is a fine-tuned version of the Qwen/Qwen3-1.7B-Base architecture, featuring approximately 1.7 billion parameters and a context length of 32768 tokens. It was developed by Kazuki1450 and trained using the TRL framework.
Key Differentiator: GRPO Training
The primary distinction of this model lies in its training methodology. It utilizes GRPO (Gradient Regularized Policy Optimization), a method introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This indicates a specific focus on enhancing the model's capabilities in mathematical reasoning.
Use Cases
- Mathematical Problem Solving: Ideal for tasks requiring robust mathematical reasoning, such as solving equations, logical deductions, or complex calculations.
- Research and Development: Can serve as a base for further fine-tuning on specialized mathematical or scientific datasets.
Training Details
The model's training procedure is publicly logged and can be visualized via Weights & Biases. It leverages specific versions of popular frameworks:
- TRL: 0.29.0
- Transformers: 4.57.6
- Pytorch: 2.9.0
- Datasets: 4.8.2
- Tokenizers: 0.22.2