cjiao/goldengoose-divsweep_goose_n128_grouporc_tau0.10-25grp
The cjiao/goldengoose-divsweep_goose_n128_grouporc_tau0.10-25grp model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is particularly suited for tasks requiring improved mathematical problem-solving and logical deduction, leveraging its 32K context length.
Loading preview...
Model Overview
This model, goldengoose-divsweep_goose_n128_grouporc_tau0.10-25grp, is a fine-tuned version of the Qwen/Qwen2.5-1.5B-Instruct base model, developed by cjiao. It features 1.5 billion parameters and supports a context length of 32,768 tokens.
Key Differentiator: GRPO Training
The primary distinction of this model lies in its training methodology. It was fine-tuned using the GRPO (Grouped Reinforcement Learning with Policy Optimization) method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This technique is specifically designed to enhance a model's mathematical reasoning abilities.
Use Cases
- Mathematical Problem Solving: Ideal for applications requiring robust mathematical reasoning and logical deduction.
- Instruction Following: Benefits from its instruction-tuned base, making it suitable for various prompt-based tasks.
Technical Details
The model was trained using the TRL library (version 0.19.1) with Transformers 4.57.6 and PyTorch 2.5.1. The training process is logged and can be visualized via Weights & Biases.