hanshan1988/wordle-grpo-Qwen3-1.7B
The hanshan1988/wordle-grpo-Qwen3-1.7B is a 2 billion parameter language model, fine-tuned from Qwen/Qwen3-1.7B. This model was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is optimized for tasks requiring advanced reasoning, leveraging its Qwen3 base and specialized training approach.
Loading preview...
Model Overview
The hanshan1988/wordle-grpo-Qwen3-1.7B is a 2 billion parameter language model, fine-tuned from the base Qwen/Qwen3-1.7B architecture. This model distinguishes itself through its specialized training methodology, utilizing GRPO (Gradient-based Reward Policy Optimization). GRPO is a technique introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models," indicating a focus on improving complex reasoning abilities.
Key Capabilities
- Enhanced Reasoning: Trained with GRPO, this model is specifically adapted to improve performance on tasks that require advanced logical and mathematical reasoning.
- Qwen3 Base: Benefits from the robust architecture and pre-training of the Qwen3-1.7B model.
- Fine-tuned Performance: Leverages the TRL (Transformers Reinforcement Learning) framework for its fine-tuning process, suggesting an optimization for specific task performance.
Training Details
The model's training procedure involved the GRPO method, as detailed in the DeepSeekMath paper. The fine-tuning was conducted using TRL version 1.7.1, with Transformers 5.12.1 and Pytorch 2.11.0.
Good For
- Applications requiring improved mathematical and logical reasoning.
- Developers looking for a Qwen3-based model with specialized reasoning enhancements.
- Experimentation with models trained using advanced reinforcement learning techniques like GRPO.