cjiao/goldengoose-divsweep_goose_n512_grouporc_tau1.00-7grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 16, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweep_goose_n512_grouporc_tau1.00-7grp model is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, this model utilizes the GRPO (Grouped Reinforcement Learning with Policy Optimization) method, as introduced in DeepSeekMath, to enhance its capabilities. It is specifically optimized for tasks requiring advanced reasoning, particularly in mathematical contexts, leveraging its 32K context length. This model is suitable for applications demanding robust logical and mathematical problem-solving from a compact LLM.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweep_goose_n512_grouporc_tau1.00-7grp is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. It was developed by cjiao and trained using the TRL framework.

Key Capabilities & Training

This model's primary differentiator lies in its training methodology. It incorporates GRPO (Grouped Reinforcement Learning with Policy Optimization), a technique detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This suggests an optimization for tasks that benefit from advanced reasoning and mathematical problem-solving.

Use Cases

Given its fine-tuning with GRPO, this model is particularly well-suited for:

  • Mathematical reasoning tasks: Applications requiring logical deduction and problem-solving in mathematical domains.
  • Instruction following: Leveraging its base as an instruction-tuned model.
  • Resource-constrained environments: As a 1.5B parameter model, it offers a balance of capability and efficiency, especially with its 32K context length.