cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau1.00_n25

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 21, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau1.00_n25 model is a 1.5 billion parameter language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct by cjiao. It was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring advanced reasoning, particularly in mathematical contexts, building upon the strong base of the Qwen2.5 architecture.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau1.00_n25 model is a 1.5 billion parameter language model, fine-tuned by cjiao from the base model Qwen/Qwen2.5-1.5B-Instruct. It leverages the TRL (Transformer Reinforcement Learning) framework for its training process.

Key Differentiator: GRPO Training

A significant aspect of this model's development is its training with GRPO (Grouped Reinforcement Learning with Policy Optimization). This method, introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300), is specifically designed to enhance a model's capabilities in mathematical reasoning. This suggests the model is optimized for complex problem-solving and logical deduction.

Capabilities

  • Enhanced Mathematical Reasoning: Through the application of the GRPO training method, this model is expected to perform well on tasks that require advanced mathematical understanding and problem-solving.
  • Instruction Following: As a fine-tuned version of an instruct model, it retains strong instruction-following capabilities, making it suitable for conversational agents and task-oriented applications.
  • Text Generation: Capable of generating coherent and contextually relevant text, as demonstrated by the quick start example.

When to Use This Model

This model is particularly well-suited for use cases where:

  • Mathematical problem-solving is a primary requirement.
  • You need a compact yet capable model (1.5B parameters) for reasoning tasks.
  • Instruction-tuned performance is crucial for generating specific responses based on user prompts.