cjiao/goldengoose-p3fu_goose_highdiv_n128_indoc_tau0.10_seed100-25grp
The cjiao/goldengoose-p3fu_goose_highdiv_n128_indoc_tau0.10_seed100-25grp model is a fine-tuned version of Qwen/Qwen2.5-1.5B-Instruct, developed by cjiao. This model was trained using the TRL library and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is specifically optimized for tasks requiring advanced mathematical problem-solving, building upon the foundation of the Qwen2.5-1.5B architecture.
Loading preview...
Model Overview
This model, cjiao/goldengoose-p3fu_goose_highdiv_n128_indoc_tau0.10_seed100-25grp, is a specialized fine-tuned variant of the Qwen/Qwen2.5-1.5B-Instruct base model. It was developed by cjiao and trained using the TRL library, a framework for Transformer Reinforcement Learning.
Key Capabilities
- Enhanced Mathematical Reasoning: A primary differentiator of this model is its training with GRPO (Gradient-based Reinforcement Learning with Policy Optimization), a method introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models". This suggests a focus on improving the model's ability to handle complex mathematical problems and logical deductions.
- Instruction Following: As it is fine-tuned from an "Instruct" model, it is designed to follow user instructions effectively for various text generation tasks.
Training Details
The model's training procedure leveraged GRPO, as detailed in the DeepSeekMath paper. The development environment included TRL 0.19.1, Transformers 4.57.6, Pytorch 2.5.1, Datasets 4.8.4, and Tokenizers 0.22.2.
When to Use This Model
This model is particularly suitable for use cases that require strong mathematical reasoning and problem-solving abilities, especially when building upon the Qwen2.5-1.5B-Instruct architecture. Its GRPO training suggests it may outperform general-purpose models on tasks involving numerical analysis, logical puzzles, or other mathematical challenges.