cjiao/goldengoose-divsweep_goose_n128_grouporc_tau2.00-25grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 15, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweep_goose_n128_grouporc_tau2.00-25grp model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, this model utilizes the GRPO method, as introduced in the DeepSeekMath paper, to enhance mathematical reasoning capabilities. With a context length of 32768 tokens, it is optimized for tasks requiring robust logical and mathematical problem-solving.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweep_goose_n128_grouporc_tau2.00-25grp is a 1.5 billion parameter instruction-tuned language model, building upon the Qwen/Qwen2.5-1.5B-Instruct base. It was fine-tuned using the TRL library.

Key Differentiator: GRPO Training

This model's unique characteristic is its training methodology, which incorporates GRPO (Grouped Reinforcement Learning with Policy Optimization). This method, detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models", is designed to significantly improve mathematical reasoning abilities in language models. This makes the model particularly adept at tasks requiring logical deduction and problem-solving.

Technical Specifications

  • Base Model: Qwen2.5-1.5B-Instruct
  • Parameters: 1.5 Billion
  • Context Length: 32768 tokens
  • Training Frameworks: TRL (0.19.1), Transformers (4.57.6), Pytorch (2.5.1), Datasets (4.8.4), Tokenizers (0.22.2)

Use Cases

This model is well-suited for applications where enhanced mathematical reasoning and logical problem-solving are critical, especially within its 32K context window.