cjiao/goldengoose-divsweepv2_lowdiv_goose_n512_indorc_tau0.10_n7
The cjiao/goldengoose-divsweepv2_lowdiv_goose_n512_indorc_tau0.10_n7 model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, it utilizes the GRPO training method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring improved reasoning, particularly in mathematical contexts, and supports a 32K context length.
Loading preview...
Overview
This model, goldengoose-divsweepv2_lowdiv_goose_n512_indorc_tau0.10_n7, is a fine-tuned variant of the Qwen/Qwen2.5-1.5B-Instruct base model, developed by cjiao. It has been specifically trained using the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models". This training approach aims to significantly improve the model's performance in complex reasoning tasks, particularly those involving mathematics.
Key Capabilities
- Enhanced Mathematical Reasoning: Leverages the GRPO training method to improve logical and mathematical problem-solving abilities.
- Instruction Following: Built upon an instruction-tuned base model, it is designed to follow user prompts effectively.
- Context Length: Supports a substantial context window of 32,768 tokens, allowing for processing longer inputs and maintaining coherence over extended interactions.
Training Details
The model was fine-tuned using the TRL library (version 0.19.1) and other standard frameworks including Transformers (4.57.6) and PyTorch (2.5.1). The GRPO method is central to its specialized training, focusing on optimizing for reasoning performance. For more technical details on GRPO, refer to the DeepSeekMath paper.
When to Use This Model
This model is particularly suitable for applications requiring a compact yet capable language model with a focus on improved reasoning, especially in mathematical or logical problem-solving scenarios. Its 1.5 billion parameters make it efficient for deployment where larger models might be impractical, while its specialized training aims to provide a performance edge in its target domain.