cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_seed200_tau0.10_n25

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 28, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_seed200_tau0.10_n25 model is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the TRL library and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring improved mathematical problem-solving, building upon the base Qwen2.5 architecture with a 32768 token context length.

Loading preview...

Model Overview

This model, goldengoose-divsweepv2_lowdiv_goose_n128_grouporc_seed200_tau0.10_n25, is a specialized fine-tuned version of the Qwen/Qwen2.5-1.5B-Instruct base model. It features 1.5 billion parameters and supports a context length of 32768 tokens.

Key Capabilities

  • Enhanced Mathematical Reasoning: The model was trained using the GRPO method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This training approach aims to significantly improve its performance on mathematical tasks.
  • Instruction Following: Inherits strong instruction-following capabilities from its Qwen2.5-Instruct base.
  • TRL Framework: Fine-tuned using the TRL library, a framework for Transformer Reinforcement Learning.

When to Use This Model

This model is particularly well-suited for applications where robust mathematical reasoning and problem-solving are critical. Developers looking for a compact yet capable model with enhanced numerical and logical processing, especially in mathematical contexts, should consider this fine-tuned variant.