Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_tok_python_alt_1_per_10_1p0_0p0_1p0_grpo_42_rule
Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_tok_python_alt_1_per_10_1p0_0p0_1p0_grpo_42_rule is a 2 billion parameter language model fine-tuned from Qwen/Qwen3-1.7B-Base with a 32768 token context length. This model was trained using the TRL framework and specifically optimized with the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is intended for tasks requiring improved logical and mathematical problem-solving, building upon the base Qwen3 architecture.
Loading preview...
Model Overview
This model, developed by Kazuki1450, is a fine-tuned version of the Qwen/Qwen3-1.7B-Base model, featuring approximately 2 billion parameters and a 32768 token context window. It leverages the TRL (Transformers Reinforcement Learning) framework for its training procedure.
Key Training Methodology
The primary differentiator for this model is its training with GRPO (Generalized Reinforcement Learning with Policy Optimization). GRPO is a method introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This indicates a specific focus on enhancing the model's capabilities in areas related to mathematical reasoning and problem-solving.
Framework Versions
The model was trained using specific versions of key frameworks:
- TRL: 0.29.0
- Transformers: 4.57.3
- Pytorch: 2.9.0
- Datasets: 4.0.0
- Tokenizers: 0.22.1
Intended Use
Given its fine-tuning with the GRPO method, this model is particularly suited for applications that benefit from improved mathematical reasoning and logical processing, building on the foundational capabilities of the Qwen3-1.7B-Base architecture.