cjiao/goldengoose-divsweep_goose_n128_grouporc_tau1.00-25grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 15, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweep_goose_n128_grouporc_tau1.00-25grp model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the Grouped Reinforcement Learning from Policy Optimization (GRPO) method, which is designed to enhance mathematical reasoning capabilities. This model is particularly suited for tasks requiring improved logical and mathematical problem-solving, building upon the base Qwen2.5 architecture.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweep_goose_n128_grouporc_tau1.00-25grp is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. It leverages a specialized training approach to enhance its reasoning abilities.

Key Characteristics

  • Base Model: Built upon the robust Qwen2.5-1.5B-Instruct architecture.
  • Training Method: Utilizes Grouped Reinforcement Learning from Policy Optimization (GRPO), a technique introduced in the context of improving mathematical reasoning in large language models, as detailed in the DeepSeekMath paper.
  • Context Length: Supports a substantial context window of 32768 tokens.

Potential Use Cases

This model is particularly well-suited for applications where enhanced logical processing and mathematical reasoning are beneficial. Developers can consider this model for tasks that require more robust problem-solving capabilities than a standard instruction-tuned model of its size, especially in domains that benefit from improved mathematical understanding.