cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_seed200_tau0.10_n25
The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_seed200_tau0.10_n25 model is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the TRL library and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring improved mathematical problem-solving, building upon the base Qwen2.5 architecture with a 32768 token context length.
Loading preview...
Model Overview
This model, goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_seed200_tau0.10_n25, is a specialized fine-tuned version of the Qwen/Qwen2.5-1.5B-Instruct base model. It features 1.5 billion parameters and supports a context length of 32768 tokens.
Key Capabilities
- Enhanced Mathematical Reasoning: The model was trained using the GRPO method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This training approach aims to significantly improve its performance on mathematical tasks.
- Instruction Following: Inherits strong instruction-following capabilities from its Qwen2.5-Instruct base.
- TRL Framework: Fine-tuned using the TRL library, a framework for Transformer Reinforcement Learning.
When to Use This Model
This model is particularly well-suited for applications where robust mathematical reasoning and problem-solving are critical. Developers looking for a compact yet capable model with enhanced numerical and logical processing, especially in mathematical contexts, should consider this fine-tuned variant.