cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau0.50_n25
The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau0.50_n25 model is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct, featuring a 32768-token context length. Developed by cjiao, this model was trained using the TRL framework and incorporates the Grouped Reinforcement Learning from Policy Optimization (GRPO) method. It is specifically optimized for enhanced reasoning capabilities, particularly in mathematical contexts, leveraging techniques from DeepSeekMath research.
Loading preview...
Model Overview
The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau0.50_n25 is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. It supports a substantial context length of 32768 tokens.
Key Differentiator: GRPO Training
This model's primary distinction lies in its training methodology. It was fine-tuned using the TRL framework and specifically incorporates Grouped Reinforcement Learning from Policy Optimization (GRPO). This method, introduced in the research behind DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, aims to significantly enhance the model's reasoning abilities.
Intended Use Cases
Given its GRPO-based training, this model is particularly well-suited for:
- Mathematical reasoning tasks: Excelling in problems requiring logical deduction and numerical understanding.
- Complex problem-solving: Benefiting from the enhanced reasoning capabilities derived from its training approach.
- Instruction-following scenarios: Building upon its
Qwen2.5-1.5B-Instructbase for general conversational and task-oriented interactions.