hkr04/qwen3-4b-grpo-dapo17k
hkr04/qwen3-4b-grpo-dapo17k is a 4 billion parameter Qwen3-based language model. It has been fine-tuned using the GRPO method on the DAPO-Math-17k dataset. This model is specifically optimized for mathematical reasoning and problem-solving tasks, leveraging its training on a specialized math dataset.
Loading preview...
Model Overview
hkr04/qwen3-4b-grpo-dapo17k is a 4 billion parameter language model built upon the Qwen3 architecture. Its primary distinction lies in its specialized training regimen: it has been fine-tuned using the GRPO (Grouped Reinforcement Learning from Human Feedback with Policy Optimization) method. This process was applied to the DAPO-Math-17k dataset, indicating a strong focus on enhancing mathematical reasoning and problem-solving capabilities.
Key Training Details
- Training Method: GRPO (Grouped Reinforcement Learning from Human Feedback with Policy Optimization)
- Dataset: DAPO-Math-17k, a specialized dataset for mathematical tasks.
- Batch Size: 32
- Group Size: 8
- Epochs: 1
- Steps: 559
- Maximum Response Length: 8192 tokens
Good For
- Mathematical Reasoning: Excels in tasks requiring logical deduction and problem-solving within a mathematical context.
- Specialized Math Applications: Suitable for applications where strong performance on mathematical datasets is crucial.
- Research in RLHF for Math: Provides a base for further experimentation with GRPO and similar fine-tuning techniques on mathematical domains.