cjiao/goldengoose-p3_goose_lowdiv_n128_indoc_tau0.10-25grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 9, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-p3_goose_lowdiv_n128_indoc_tau0.10-25grp model is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the GRPO method, as introduced in the DeepSeekMath paper, which focuses on enhancing mathematical reasoning. This model is optimized for tasks requiring improved reasoning capabilities, particularly in areas where GRPO's methodology is beneficial.

Loading preview...

Model Overview

This model, goldengoose-p3_goose_lowdiv_n128_indoc_tau0.10-25grp, is a 1.5 billion parameter language model derived from the Qwen/Qwen2.5-1.5B-Instruct base model. It has been specifically fine-tuned using the TRL (Transformer Reinforcement Learning) framework.

Key Differentiator: GRPO Training

A core aspect of this model's development is its training with the GRPO (Gradient-based Reward Policy Optimization) method. This technique, detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models," aims to significantly improve a model's mathematical reasoning abilities. By applying GRPO, this model is designed to exhibit enhanced performance in tasks that benefit from advanced reasoning.

Training Details

The fine-tuning process leveraged TRL version 0.19.1, with Transformers 4.57.6, Pytorch 2.5.1, Datasets 4.8.4, and Tokenizers 0.22.2. The training run can be visualized via Weights & Biases, indicating a structured and monitored development process.

Use Cases

This model is particularly suitable for applications requiring robust reasoning, especially in domains that can benefit from the GRPO training methodology. Developers can integrate it using the Hugging Face pipeline for text generation tasks.