swagger00/Qwen3-8B-GRPO-GSM8K
Qwen3-8B-GRPO-GSM8K is an 8 billion parameter Qwen3 model developed by swagger00, fine-tuned using GRPO and LoRA. This model is specifically optimized for mathematical reasoning tasks, demonstrating a significant improvement in GSM8K accuracy. It is designed for applications requiring robust performance in quantitative problem-solving.
Loading preview...
Model Overview
Qwen3-8B-GRPO-GSM8K is an 8 billion parameter Qwen3 model, developed by swagger00, that has been fine-tuned using the GRPO (Gradient Regularized Policy Optimization) method combined with LoRA (Low-Rank Adaptation). The LoRA weights have been merged directly into the model for seamless use. This model is particularly notable for its strong performance on mathematical reasoning benchmarks.
Key Capabilities
- Enhanced Mathematical Reasoning: The model shows significant improvement in solving mathematical word problems, specifically on the GSM8K dataset.
- GRPO and LoRA Fine-tuning: Utilizes advanced training techniques (GRPO and LoRA with rank 64, alpha 32) to achieve specialized performance.
- Optimized Checkpoint: The final model checkpoint (global_step_116) was selected based on achieving the highest validation accuracy during training.
Performance Highlights
During training, the model's accuracy on the GSM8K validation set improved substantially:
- Initial Model: 24.94% accuracy
- After 116 Steps: 72.33% accuracy
These evaluation scores were recorded under the training recipe's settings, using the GSM8K test split for validation. The model is well-suited for applications requiring accurate quantitative problem-solving capabilities.
Good For
- Mathematical problem-solving and reasoning tasks.
- Educational tools requiring arithmetic and logic.
- Applications where robust performance on benchmarks like GSM8K is critical.