cjiao/goldengoose-divsweep_goose_n128_grouporc_tau2.00-25grp
The cjiao/goldengoose-divsweep_goose_n128_grouporc_tau2.00-25grp model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, this model utilizes the GRPO method, as introduced in the DeepSeekMath paper, to enhance mathematical reasoning capabilities. With a context length of 32768 tokens, it is optimized for tasks requiring robust logical and mathematical problem-solving.
Loading preview...
Model Overview
The cjiao/goldengoose-divsweep_goose_n128_grouporc_tau2.00-25grp is a 1.5 billion parameter instruction-tuned language model, building upon the Qwen/Qwen2.5-1.5B-Instruct base. It was fine-tuned using the TRL library.
Key Differentiator: GRPO Training
This model's unique characteristic is its training methodology, which incorporates GRPO (Grouped Reinforcement Learning with Policy Optimization). This method, detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models", is designed to significantly improve mathematical reasoning abilities in language models. This makes the model particularly adept at tasks requiring logical deduction and problem-solving.
Technical Specifications
- Base Model: Qwen2.5-1.5B-Instruct
- Parameters: 1.5 Billion
- Context Length: 32768 tokens
- Training Frameworks: TRL (0.19.1), Transformers (4.57.6), Pytorch (2.5.1), Datasets (4.8.4), Tokenizers (0.22.2)
Use Cases
This model is well-suited for applications where enhanced mathematical reasoning and logical problem-solving are critical, especially within its 32K context window.