cjiao/goldengoose-p3_goose_lowdiv_n128_indoc_tau1.00-25grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 9, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-p3_goose_lowdiv_n128_indoc_tau1.00-25grp model is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, this model was trained using the GRPO method, which is designed to enhance mathematical reasoning in open language models. It is optimized for tasks requiring improved reasoning capabilities, particularly in mathematical contexts, leveraging its Qwen2.5 base architecture.

Loading preview...

Model Overview

The cjiao/goldengoose-p3_goose_lowdiv_n128_indoc_tau1.00-25grp is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. This model was developed by cjiao and trained using the TRL framework.

Key Training Details

A significant aspect of this model's development is its training methodology. It utilizes GRPO (Grouped Reinforcement Learning with Policy Optimization), a method introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This indicates a specific focus on enhancing the model's ability to handle complex reasoning tasks, particularly within mathematical domains.

Intended Use Cases

Given its fine-tuning with the GRPO method, this model is particularly well-suited for applications that require:

  • Mathematical Reasoning: Tasks involving problem-solving, logical deduction, and numerical analysis.
  • Instruction Following: Leveraging its Instruct base, it can effectively follow user prompts and generate relevant responses.
  • General Language Generation: As a Qwen2.5-based model, it retains strong capabilities for various text generation tasks.