cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_indorc_tau1.00_n25
TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 21, 2026Architecture:Transformer Featherless Exclusive Cold
The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_indorc_tau1.00_n25 is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring robust reasoning, particularly in mathematical contexts, and supports a 32K context length.
Loading preview...
Model Overview
This model, goldengoose-divsweepv2_lowdiv_goose_n128_indorc_tau1.00_n25, is a specialized instruction-tuned variant of the Qwen/Qwen2.5-1.5B-Instruct base model. It features 1.5 billion parameters and supports a substantial context length of 32,768 tokens.
Key Capabilities & Training
- Enhanced Reasoning: The model was fine-tuned using GRPO (Gradient-based Reward Policy Optimization), a method introduced in the DeepSeekMath paper. This training approach specifically aims to improve the model's performance in complex reasoning tasks, particularly those involving mathematics.
- Instruction Following: As an instruction-tuned model, it is designed to accurately follow user prompts and generate relevant responses.
- Frameworks: The training process leveraged the TRL library for transformer reinforcement learning.
When to Use This Model
- Mathematical Reasoning: Ideal for applications requiring strong logical and mathematical problem-solving abilities.
- Instruction-Based Tasks: Suitable for general instruction-following scenarios where a compact yet capable model is needed.
- Resource-Efficient Deployment: Its 1.5B parameter size makes it a good choice for environments with computational constraints, while still offering advanced reasoning features.