Palind/qwen25-0.5b-ppo-gsm8k
Palind/qwen25-0.5b-ppo-gsm8k is a 0.5 billion parameter Qwen2.5-Instruct model fine-tuned using PPO on the GSM8K dataset, specializing in mathematical reasoning and arithmetic problem-solving. Developed by Palind, this model demonstrates significantly enhanced performance on GSM8K-style arithmetic tasks compared to its base model, achieving 55.57% strict accuracy. It is optimized for generating step-by-step solutions and numerical answers for math problems within a 32768 token context length.
Loading preview...
Overview
Palind/qwen25-0.5b-ppo-gsm8k is an experimental 0.5 billion parameter model based on Qwen/Qwen2.5-0.5B-Instruct, fine-tuned using Proximal Policy Optimization (PPO) on the GSM8K dataset. This model is specifically designed to improve performance on grade school mathematical reasoning problems. The training utilized the verl framework, focusing on generating accurate numerical answers in a specific format.
Key Capabilities
- Enhanced Mathematical Reasoning: Significantly improves strict accuracy on GSM8K arithmetic problems from 0.61% (base model) to 55.57% after PPO training.
- Structured Output: Optimized to provide final answers in the
#### numberformat, aligning with standard GSM8K evaluation. - PPO Fine-tuning: Demonstrates the effectiveness of PPO with GAE advantages and a rule-based reward system for domain-specific task improvement.
Limitations and Considerations
- Domain Specificity: Primarily trained and evaluated on GSM8K-style arithmetic; general capabilities may degrade in other domains.
- Experimental Nature: This is a small research artifact, not intended for production, and results are from a single experimental run.
- Sensitivity: Performance is sensitive to prompt format, sampling settings, and the reward parser used during training.