cjiao/goldengoose-p3_goose_highdiv_n128_random-25grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 9, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-p3_goose_highdiv_n128_random-25grp model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, it leverages the GRPO method, introduced in DeepSeekMath, for enhanced training. This model is optimized for tasks requiring advanced reasoning, particularly in mathematical contexts, and supports a 32768 token context length.

Loading preview...

Model Overview

cjiao/goldengoose-p3_goose_highdiv_n128_random-25grp is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. It was developed by cjiao and trained using the TRL framework.

Key Capabilities

  • Enhanced Reasoning: This model incorporates the GRPO (Gradient-based Reward Policy Optimization) method, originally introduced in the DeepSeekMath paper, to improve its reasoning capabilities.
  • Instruction Following: As a fine-tuned instruction model, it is designed to follow user prompts effectively.
  • Extended Context Window: Supports a context length of 32768 tokens, allowing for processing longer inputs and generating more coherent, extended responses.

Training Details

The model's training procedure utilized GRPO, a method known for pushing the limits of mathematical reasoning in open language models. This suggests a focus on improving logical and analytical processing.

Good For

  • Mathematical Reasoning Tasks: Given its training with the GRPO method from DeepSeekMath, this model is particularly well-suited for tasks that require strong mathematical or logical reasoning.
  • Instruction-based Applications: Ideal for applications where the model needs to accurately interpret and respond to specific instructions.
  • Long Context Processing: Beneficial for use cases requiring the model to understand and generate text based on extensive contextual information.