cjiao/goldengoose-divsweep_goose_n128_indorc_tau0.10_seed200-25grp
TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 26, 2026Architecture:Transformer Featherless Exclusive Cold
The cjiao/goldengoose-divsweep_goose_n128_indorc_tau0.10_seed200-25grp model is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, this model was trained using the GRPO method, which is specifically designed to enhance mathematical reasoning capabilities. It is optimized for tasks requiring robust logical and mathematical problem-solving, making it suitable for applications where precise reasoning is critical.
Loading preview...
Model Overview
The cjiao/goldengoose-divsweep_goose_n128_indorc_tau0.10_seed200-25grp is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. It leverages a context length of 32768 tokens, providing ample capacity for complex inputs.
Key Capabilities
- Enhanced Mathematical Reasoning: This model was trained using the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models". This specialized training focuses on improving the model's ability to handle mathematical and logical problems.
- Instruction Following: As a fine-tuned version of an instruction-tuned model, it is designed to follow user prompts effectively, making it suitable for interactive applications.
- TRL Framework: The fine-tuning process utilized the TRL (Transformer Reinforcement Learning) library, indicating a reinforcement learning approach was applied to optimize its performance.
Good For
- Mathematical Problem Solving: Its GRPO training makes it particularly well-suited for tasks that require strong mathematical reasoning and logical deduction.
- Instruction-based Applications: Ideal for chatbots, virtual assistants, or any application where the model needs to respond accurately to specific instructions.
- Research in RLHF/RLFT: Provides a practical example of a model fine-tuned with advanced reinforcement learning techniques like GRPO, useful for researchers exploring these methods.