Kazuki1450/Qwen3-1.7B-Base_dsum_3_6_tok_python_1p0_0p0_1p0_grpo_sapo_42_rule
This model is a 2 billion parameter language model, fine-tuned from Qwen3-1.7B-Base by Kazuki1450. It utilizes the GRPO training method, as detailed in the DeepSeekMath paper, which focuses on mathematical reasoning. This fine-tuning aims to enhance the model's capabilities, particularly in areas related to reasoning and problem-solving, making it suitable for tasks requiring structured logical thought.
Loading preview...
Model Overview
This model, developed by Kazuki1450, is a fine-tuned version of the Qwen3-1.7B-Base architecture, featuring approximately 2 billion parameters and a 32768 token context length. It was trained using the TRL framework and specifically incorporates the GRPO (Gradient-based Reward Policy Optimization) method.
Key Capabilities
- Enhanced Reasoning: The integration of the GRPO method, derived from the DeepSeekMath research, suggests a focus on improving the model's ability to handle complex reasoning tasks.
- Fine-tuned Performance: As a fine-tuned variant, it is expected to offer specialized performance beyond the base Qwen3-1.7B model, particularly in areas where GRPO provides benefits.
Training Methodology
The model's training procedure highlights the use of GRPO, a technique introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models". This indicates an optimization strategy aimed at improving logical and mathematical problem-solving capabilities. The training utilized standard frameworks including TRL, Transformers, Pytorch, Datasets, and Tokenizers.
Good For
- Applications requiring improved reasoning and logical inference.
- Tasks that could benefit from a model trained with methods designed for mathematical problem-solving.
- Developers looking for a specialized Qwen3-1.7B variant with a focus on structured output and coherence.