cjiao/goldengoose-p3fu_goose_highdiv_n128_indoc_tau0.10_seed100-25grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 10, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-p3fu_goose_highdiv_n128_indoc_tau0.10_seed100-25grp model is a fine-tuned version of Qwen/Qwen2.5-1.5B-Instruct, developed by cjiao. This model was trained using the TRL library and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is specifically optimized for tasks requiring advanced mathematical problem-solving, building upon the foundation of the Qwen2.5-1.5B architecture.

Loading preview...

Model Overview

This model, cjiao/goldengoose-p3fu_goose_highdiv_n128_indoc_tau0.10_seed100-25grp, is a specialized fine-tuned variant of the Qwen/Qwen2.5-1.5B-Instruct base model. It was developed by cjiao and trained using the TRL library, a framework for Transformer Reinforcement Learning.

Key Capabilities

  • Enhanced Mathematical Reasoning: A primary differentiator of this model is its training with GRPO (Gradient-based Reinforcement Learning with Policy Optimization), a method introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models". This suggests a focus on improving the model's ability to handle complex mathematical problems and logical deductions.
  • Instruction Following: As it is fine-tuned from an "Instruct" model, it is designed to follow user instructions effectively for various text generation tasks.

Training Details

The model's training procedure leveraged GRPO, as detailed in the DeepSeekMath paper. The development environment included TRL 0.19.1, Transformers 4.57.6, Pytorch 2.5.1, Datasets 4.8.4, and Tokenizers 0.22.2.

When to Use This Model

This model is particularly suitable for use cases that require strong mathematical reasoning and problem-solving abilities, especially when building upon the Qwen2.5-1.5B-Instruct architecture. Its GRPO training suggests it may outperform general-purpose models on tasks involving numerical analysis, logical puzzles, or other mathematical challenges.