cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau2.00_n25

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 21, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau2.00_n25 model is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, it utilizes the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is particularly suited for tasks requiring improved logical and mathematical problem-solving, building upon the base Qwen2.5 architecture with a 32K context length.

Loading preview...

Model Overview

This model, goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau2.00_n25, is a 1.5 billion parameter language model developed by cjiao. It is a fine-tuned version of the Qwen/Qwen2.5-1.5B-Instruct base model, leveraging its robust architecture and a substantial 32,768 token context length.

Key Capabilities

  • Enhanced Mathematical Reasoning: The model was trained using the GRPO method, as introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models." This training approach specifically targets and improves the model's ability to handle complex mathematical and logical reasoning tasks.
  • Instruction Following: As a fine-tuned version of an instruction-tuned model, it is designed to follow user prompts effectively.
  • TRL Framework: The fine-tuning process was conducted using the TRL (Transformer Reinforcement Learning) library, indicating a focus on optimizing model behavior through reinforcement learning techniques.

When to Use This Model

  • Mathematical Problem Solving: Ideal for applications requiring strong mathematical reasoning, logical deduction, and problem-solving, benefiting from the GRPO training.
  • General Instruction-Following: Suitable for a wide range of natural language processing tasks where clear and concise instruction adherence is crucial.
  • Research and Experimentation: Provides a specialized base for further research into mathematical reasoning and reinforcement learning applications in LLMs, particularly within the Qwen2.5 family.