cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_indorc_tau1.00_n25

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 21, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_indorc_tau1.00_n25 is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring robust reasoning, particularly in mathematical contexts, and supports a 32K context length.

Loading preview...

Model Overview

This model, goldengoose-divsweepv2_lowdiv_goose_n128_indorc_tau1.00_n25, is a specialized instruction-tuned variant of the Qwen/Qwen2.5-1.5B-Instruct base model. It features 1.5 billion parameters and supports a substantial context length of 32,768 tokens.

Key Capabilities & Training

  • Enhanced Reasoning: The model was fine-tuned using GRPO (Gradient-based Reward Policy Optimization), a method introduced in the DeepSeekMath paper. This training approach specifically aims to improve the model's performance in complex reasoning tasks, particularly those involving mathematics.
  • Instruction Following: As an instruction-tuned model, it is designed to accurately follow user prompts and generate relevant responses.
  • Frameworks: The training process leveraged the TRL library for transformer reinforcement learning.

When to Use This Model

  • Mathematical Reasoning: Ideal for applications requiring strong logical and mathematical problem-solving abilities.
  • Instruction-Based Tasks: Suitable for general instruction-following scenarios where a compact yet capable model is needed.
  • Resource-Efficient Deployment: Its 1.5B parameter size makes it a good choice for environments with computational constraints, while still offering advanced reasoning features.