cjiao/goldengoose-divsweepv2_lowdiv_goose_n512_indorc_tau2.00_n7
TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 21, 2026Architecture:Transformer Featherless Exclusive Cold
The cjiao/goldengoose-divsweepv2_lowdiv_goose_n512_indorc_tau2.00_n7 model is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, this model utilizes the GRPO training method, which is specifically designed to enhance mathematical reasoning capabilities. It is optimized for tasks requiring robust logical and mathematical problem-solving, leveraging a 32768 token context length for complex inputs.
Loading preview...
Overview
This model, goldengoose-divsweepv2_lowdiv_goose_n512_indorc_tau2.00_n7, is a 1.5 billion parameter language model fine-tuned by cjiao. It is built upon the Qwen/Qwen2.5-1.5B-Instruct architecture and leverages a substantial 32768 token context length.
Key Capabilities
- Enhanced Mathematical Reasoning: The model was trained using the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper. This technique specifically aims to improve the model's ability to handle mathematical problems and logical reasoning tasks.
- Instruction Following: As a fine-tuned version of an instruction-tuned base model, it is designed to follow user instructions effectively.
- TRL Framework: The fine-tuning process was conducted using the TRL (Transformer Reinforcement Learning) library, indicating a focus on reinforcement learning from human feedback or similar optimization strategies.
Good For
- Mathematical Problem Solving: Ideal for applications requiring accurate mathematical reasoning and problem-solving, benefiting from its GRPO-based training.
- Complex Instruction Following: Suitable for tasks where detailed and multi-step instructions need to be processed and executed.
- Research and Experimentation: Provides a compact yet capable model for exploring the effects of GRPO on smaller language models, particularly in the domain of mathematical intelligence.