dekangli/Qwen2.5-1.5B-GRPO-v5
dekangli/Qwen2.5-1.5B-GRPO-v5 is a 1.5 billion parameter language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the GRPO method on the OpenR1-Math-220k dataset, specializing it for mathematical reasoning tasks. This model is optimized to enhance performance in complex mathematical problem-solving.
Loading preview...
Model Overview
This model, dekangli/Qwen2.5-1.5B-GRPO-v5, is a specialized 1.5 billion parameter language model. It is built upon the Qwen/Qwen2.5-1.5B-Instruct architecture and has been further fine-tuned to excel in specific domains.
Key Capabilities
- Mathematical Reasoning: The model's primary strength lies in mathematical problem-solving, achieved through fine-tuning on the open-r1/OpenR1-Math-220k dataset.
- GRPO Training: It leverages the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper, to enhance its reasoning abilities.
Training Details
The model was trained using the TRL framework (version 0.18.0.dev0) with Transformers 4.52.0.dev0 and Pytorch 2.6.0. The GRPO method is detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300).
When to Use This Model
This model is particularly well-suited for applications requiring robust mathematical reasoning and problem-solving, especially where the Qwen2.5-1.5B-Instruct base model needs enhanced mathematical capabilities.