cjiao/goldengoose-divsweep_goose_n512_grouporc_tau0.10-7grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 16, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweep_goose_n512_grouporc_tau0.10-7grp model is a 1.5 billion parameter language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the GRPO method, as introduced in the DeepSeekMath paper, which focuses on mathematical reasoning. This model is optimized for tasks requiring enhanced reasoning capabilities, leveraging its specialized training approach.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweep_goose_n512_grouporc_tau0.10-7grp is a 1.5 billion parameter language model, derived from the Qwen/Qwen2.5-1.5B-Instruct base model. It has been fine-tuned using the TRL framework, incorporating a specific training methodology known as GRPO.

Key Training Details

  • Base Model: Qwen/Qwen2.5-1.5B-Instruct
  • Fine-tuning Framework: TRL (Transformer Reinforcement Learning)
  • Training Method: GRPO (Grouped Reinforcement Learning with Policy Optimization), a technique highlighted in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This suggests an optimization for tasks that benefit from advanced reasoning.
  • Context Length: The model supports a context length of 32768 tokens.

Potential Use Cases

Given its fine-tuning with GRPO, this model is likely well-suited for:

  • Tasks requiring enhanced logical and mathematical reasoning.
  • Applications where the base Qwen2.5-1.5B-Instruct model's capabilities are further specialized for complex problem-solving.
  • Exploration of models trained with advanced reinforcement learning techniques for improved performance in specific domains.