Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_tok_python_alt_1_per_5_1p0_0p0_1p0_grpo_42_rule
Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_tok_python_alt_1_per_5_1p0_0p0_1p0_grpo_42_rule is a 2 billion parameter language model fine-tuned from Qwen/Qwen3-1.7B-Base. This model was trained using the GRPO (Grouped Reinforcement Learning with Policy Optimization) method, as introduced in the DeepSeekMath paper, to enhance specific reasoning capabilities. It is optimized for tasks that benefit from advanced mathematical reasoning and structured problem-solving approaches. The model has a context length of 32768 tokens, making it suitable for processing extensive inputs.
Loading preview...
Model Overview
This model, Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_tok_python_alt_1_per_5_1p0_0p0_1p0_grpo_42_rule, is a fine-tuned version of the Qwen/Qwen3-1.7B-Base architecture, featuring approximately 2 billion parameters and a 32K token context window. It was developed by Kazuki1450 and trained using the TRL library.
Key Differentiator: GRPO Training
The primary distinction of this model lies in its training methodology. It utilizes GRPO (Grouped Reinforcement Learning with Policy Optimization), a technique detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This method is designed to improve the model's ability to handle complex reasoning tasks, particularly those involving mathematical or logical structures.
Potential Use Cases
- Mathematical Reasoning: Ideal for applications requiring robust mathematical problem-solving.
- Structured Problem Solving: Suitable for tasks that benefit from a model trained with advanced optimization techniques for reasoning.
- Research and Experimentation: Provides a base for further research into GRPO-enhanced models and their performance on specific reasoning benchmarks.
Training Details
The model was fine-tuned using the TRL (Transformers Reinforcement Learning) framework, with specific versions of TRL (0.29.0), Transformers (4.57.3), Pytorch (2.9.0), Datasets (4.0.0), and Tokenizers (0.22.1) used during its development.