cjiao/goldengoose-divsweep_goose_n512_grouporc_tau2.00-7grp
The cjiao/goldengoose-divsweep_goose_n512_grouporc_tau2.00-7grp model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, this model utilizes the GRPO method, originally introduced for mathematical reasoning in DeepSeekMath. It is specifically optimized for enhanced reasoning capabilities, making it suitable for tasks requiring structured thought processes over its 32K context length.
Loading preview...
Model Overview
The cjiao/goldengoose-divsweep_goose_n512_grouporc_tau2.00-7grp is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. This model was developed by cjiao and leverages the GRPO (Grouped Reasoning Policy Optimization) method, a technique highlighted in the DeepSeekMath paper for improving mathematical reasoning in large language models. The fine-tuning process was conducted using the TRL (Transformer Reinforcement Learning) framework.
Key Capabilities
- Enhanced Reasoning: Incorporates the GRPO method to improve the model's ability to handle complex reasoning tasks, drawing inspiration from its application in mathematical problem-solving.
- Instruction Following: As a fine-tuned version of an instruction-tuned model, it is designed to follow user instructions effectively.
- Context Handling: Supports a substantial context length of 32,768 tokens, allowing for processing and generating longer, more detailed responses.
Training Details
The model's training procedure utilized specific versions of popular frameworks, including TRL 0.19.1, Transformers 4.57.6, PyTorch 2.5.1, Datasets 4.8.4, and Tokenizers 0.22.2. The use of GRPO suggests a focus on improving the model's internal reasoning mechanisms beyond standard instruction tuning.
Good For
This model is particularly well-suited for applications requiring robust reasoning and structured problem-solving, especially where the GRPO method's benefits in mathematical or logical tasks could be advantageous. Its 1.5B parameter count makes it a relatively efficient option for such tasks, balancing performance with computational resources.