cjiao/goldengoose-divsweep_goose_n128_random-25grp
The cjiao/goldengoose-divsweep_goose_n128_random-25grp is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct with a 32K context length. This model was trained using the GRPO method, which is specifically designed to enhance mathematical reasoning capabilities. It is optimized for tasks requiring advanced logical and mathematical problem-solving, building upon the robust Qwen2.5 architecture.
Loading preview...
Model Overview
The cjiao/goldengoose-divsweep_goose_n128_random-25grp is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. It leverages a substantial 32,768 token context window, making it suitable for processing longer inputs and generating extended responses.
Key Capabilities
- Enhanced Mathematical Reasoning: This model was specifically trained using the GRPO (Gradient-based Reasoning Policy Optimization) method, as introduced in the DeepSeekMath research. This training approach aims to significantly improve the model's ability to handle complex mathematical and logical reasoning tasks.
- Instruction Following: As a fine-tuned instruction model, it is designed to accurately interpret and execute user prompts, providing relevant and coherent responses.
- Qwen2.5 Architecture: Built upon the robust Qwen2.5 series, it inherits a strong foundation for general language understanding and generation.
Training Details
The model's fine-tuning process utilized the TRL library (version 0.19.1) and incorporated the GRPO method. This method is detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300).
Good For
- Applications requiring strong mathematical problem-solving.
- Tasks that benefit from advanced logical reasoning.
- Instruction-following scenarios where precise and contextually relevant outputs are crucial.