cjiao/goldengoose-p3fu_goose_highdiv_n128_grpoc_tau0.10_seed200-25grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 10, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-p3fu_goose_highdiv_n128_grpoc_tau0.10_seed200-25grp model is a 1.5 billion parameter language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring advanced mathematical problem-solving and logical deduction, leveraging a 32768-token context length.

Loading preview...

Model Overview

The cjiao/goldengoose-p3fu_goose_highdiv_n128_grpoc_tau0.10_seed200-25grp is a 1.5 billion parameter language model, building upon the base architecture of Qwen/Qwen2.5-1.5B-Instruct. It has been specifically fine-tuned using the TRL (Transformer Reinforcement Learning) framework.

Key Capabilities & Training

A significant differentiator for this model is its training methodology, which incorporates GRPO (Gradient-based Reinforcement Learning with Policy Optimization). This method, introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models," aims to significantly enhance the model's ability to perform complex mathematical reasoning tasks. The model leverages a substantial context length of 32768 tokens, which is beneficial for processing and understanding intricate problems.

Use Cases

Given its specialized training with GRPO, this model is particularly well-suited for applications requiring:

  • Mathematical problem-solving: Excelling in tasks that demand logical deduction and numerical accuracy.
  • Reasoning-intensive queries: Handling complex questions where understanding underlying principles is crucial.
  • Instruction-following in technical domains: Benefiting from its instruction-tuned base and specialized fine-tuning.

Developers looking for a compact yet powerful model with enhanced mathematical and reasoning capabilities should consider this fine-tuned variant.