cjiao/goldengoose-p3fu_goose_highdiv_n128_grpoc_tau0.10_seed100-25grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 10, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-p3fu_goose_highdiv_n128_grpoc_tau0.10_seed100-25grp model is a fine-tuned version of Qwen/Qwen2.5-1.5B-Instruct, developed by cjiao. It was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning in language models. This model is particularly suited for tasks requiring improved reasoning capabilities, building upon the base Qwen2.5-1.5B-Instruct architecture.

Loading preview...

Model Overview

This model, cjiao/goldengoose-p3fu_goose_highdiv_n128_grpoc_tau0.10_seed100-25grp, is a specialized fine-tune of the Qwen/Qwen2.5-1.5B-Instruct base model. It was developed by cjiao and trained using the TRL (Transformer Reinforcement Learning) framework.

Key Differentiator: GRPO Training

A significant aspect of this model's training is the application of GRPO (Gradient-based Reinforcement Learning with Policy Optimization). This method, introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300), aims to enhance the mathematical reasoning capabilities of large language models. By leveraging GRPO, this model is expected to exhibit improved performance in tasks that require logical and mathematical problem-solving.

Training Environment

The model's training procedure utilized specific versions of popular machine learning frameworks:

  • TRL: 0.19.1
  • Transformers: 4.57.6
  • Pytorch: 2.5.1
  • Datasets: 4.8.4
  • Tokenizers: 0.22.2

Use Cases

Given its fine-tuning with the GRPO method, this model is particularly well-suited for applications where enhanced reasoning, especially in mathematical or logical contexts, is beneficial. Developers can integrate it into systems requiring more robust problem-solving abilities than a standard instruction-tuned model of its size might offer.