cjiao/goldengoose-divsweep_goose_n512_random_seed200-7grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 27, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweep_goose_n512_random_seed200-7grp model is a 1.5 billion parameter language model, fine-tuned by cjiao from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is particularly suited for tasks requiring advanced mathematical problem-solving and logical deduction.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweep_goose_n512_random_seed200-7grp is a 1.5 billion parameter language model, fine-tuned from the base Qwen/Qwen2.5-1.5B-Instruct architecture. Developed by cjiao, this model leverages the TRL (Transformer Reinforcement Learning) framework for its training process.

Key Differentiator: GRPO Training

A significant aspect of this model's development is its training with GRPO (Gradient-based Reward Policy Optimization). This method, introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models," is specifically designed to improve a model's proficiency in mathematical reasoning tasks. This makes goldengoose-divsweep_goose_n512_random_seed200-7grp distinct from general-purpose instruction-tuned models.

Capabilities

  • Enhanced Mathematical Reasoning: Optimized through the GRPO method, the model is expected to perform well on tasks requiring mathematical problem-solving and logical deduction.
  • Instruction Following: As a fine-tuned version of an instruction-tuned model, it retains strong capabilities in understanding and executing user instructions.

When to Use This Model

This model is particularly well-suited for applications where:

  • Mathematical problem-solving is a primary requirement.
  • You need a relatively compact model (1.5B parameters) with specialized reasoning abilities.
  • You are exploring the impact of GRPO on language model performance for mathematical tasks.