cjiao/goldengoose-divsweep_goose_n128_grouporc_tau0.10_seed200-25grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 26, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweep_goose_n128_grouporc_tau0.10_seed200-25grp model is a 1.5 billion parameter language model fine-tuned by cjiao, based on the Qwen2.5-1.5B-Instruct architecture. This model was specifically trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is optimized for tasks requiring robust logical and mathematical problem-solving, making it suitable for applications where precise reasoning is critical.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweep_goose_n128_grouporc_tau0.10_seed200-25grp is a 1.5 billion parameter language model, fine-tuned by cjiao from the Qwen2.5-1.5B-Instruct base model. It leverages the TRL (Transformer Reinforcement Learning) framework for its training process.

Key Capabilities & Training

This model's primary differentiator is its training methodology, which incorporates GRPO (Grouped Reinforcement Learning with Policy Optimization). GRPO is a method introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This specialized training aims to significantly enhance the model's mathematical reasoning abilities.

Use Cases

Given its GRPO-enhanced training, this model is particularly well-suited for:

  • Mathematical problem-solving: Excelling in tasks that require logical deduction and numerical computation.
  • Reasoning-intensive applications: Where precise and structured thought processes are paramount.
  • Educational tools: Assisting with explanations and solutions for complex mathematical concepts.

Developers can quickly integrate this model using the Hugging Face transformers library for text generation tasks, as demonstrated in the quick start example.