cjiao/goldengoose-divsweep_goose_n512_grouporc_tau0.50-7grp
The cjiao/goldengoose-divsweep_goose_n512_grouporc_tau0.50-7grp model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring improved reasoning, particularly in mathematical contexts, leveraging its Qwen2.5 base architecture and a 32K context length.
Loading preview...
Model Overview
This model, goldengoose-divsweep_goose_n512_grouporc_tau0.50-7grp, is a 1.5 billion parameter instruction-tuned language model. It is a fine-tuned version of the Qwen/Qwen2.5-1.5B-Instruct base model, developed by cjiao. The fine-tuning process utilized the TRL (Transformer Reinforcement Learning) framework.
Key Differentiator: GRPO Training
A significant aspect of this model is its training methodology. It was fine-tuned using GRPO (Grouped Reinforcement Learning with Policy Optimization), a method introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This indicates a specific optimization for tasks that benefit from enhanced mathematical reasoning.
Capabilities and Use Cases
Given its foundation in Qwen2.5-1.5B-Instruct and specialized GRPO training, this model is likely to excel in:
- Instruction-following tasks: Inheriting capabilities from its instruction-tuned base.
- Mathematical reasoning: The GRPO training suggests improved performance on problems requiring logical and mathematical deduction.
- General text generation: For tasks where a compact yet capable model with a 32K context length is beneficial.
Technical Details
- Base Model: Qwen/Qwen2.5-1.5B-Instruct
- Training Framework: TRL (version 0.19.1)
- Training Method: GRPO, as detailed in the DeepSeekMath paper.
This model is suitable for developers looking for a smaller, efficient language model with a focus on improved reasoning, particularly in mathematical domains, building upon the strong foundation of the Qwen2.5 architecture.