cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau0.10_n25

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 21, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau0.10_n25 model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring robust mathematical problem-solving and logical deduction, leveraging its 32768 token context length.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_tau0.10_n25 is a 1.5 billion parameter instruction-tuned language model, building upon the base of Qwen/Qwen2.5-1.5B-Instruct. It was fine-tuned using the TRL library and specifically incorporates the GRPO (Grouped Reinforcement Learning with Policy Optimization) method.

Key Differentiator

The primary distinction of this model lies in its training methodology, which utilizes GRPO. This technique, introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models", is designed to significantly enhance a model's mathematical reasoning abilities and logical problem-solving.

Capabilities

  • Enhanced Mathematical Reasoning: Optimized for tasks that require complex mathematical understanding and problem-solving.
  • Instruction Following: Benefits from its instruction-tuned base, making it suitable for various prompt-based applications.
  • Extended Context: Features a 32768 token context length, allowing for processing and generating longer, more detailed responses.

Ideal Use Cases

This model is particularly well-suited for applications where strong mathematical and logical reasoning is crucial, such as:

  • Solving mathematical problems and equations.
  • Generating explanations for complex logical sequences.
  • Tasks requiring precise numerical understanding and manipulation.