cjiao/goldengoose-divsweep_goose_n128_indorc_tau1.00-25grp
The cjiao/goldengoose-divsweep_goose_n128_indorc_tau1.00-25grp model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is particularly suited for tasks requiring improved reasoning, building upon its Qwen2.5 base.
Loading preview...
Model Overview
The cjiao/goldengoose-divsweep_goose_n128_indorc_tau1.00-25grp is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. It leverages a context length of 32768 tokens, making it suitable for processing longer inputs.
Key Training Details
This model was developed using the TRL framework. A significant aspect of its training methodology is the application of GRPO (Gradient-based Reward Policy Optimization). GRPO is a method introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models," indicating a focus on enhancing the model's reasoning abilities, particularly in mathematical contexts.
Potential Use Cases
Given its foundation in Qwen2.5-1.5B-Instruct and the application of GRPO, this model is likely to perform well in:
- Instruction-following tasks: Benefiting from its instruction-tuned base.
- Reasoning-intensive applications: Especially those involving mathematical or logical problem-solving, due to the GRPO training.
- Applications requiring a balance of performance and efficiency: Its 1.5B parameter count offers a lighter footprint compared to larger models while still providing enhanced reasoning capabilities.