cjiao/goldengoose-p3fu_goose_highdiv_n128_random_seed100-25grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 10, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-p3fu_goose_highdiv_n128_random_seed100-25grp model is a 1.5 billion parameter language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities in open language models. This model is optimized for tasks requiring advanced reasoning, particularly in mathematical contexts, leveraging its 32768 token context length.

Loading preview...

Model Overview

The cjiao/goldengoose-p3fu_goose_highdiv_n128_random_seed100-25grp is a 1.5 billion parameter instruction-tuned language model, built upon the Qwen2.5-1.5B-Instruct architecture. This model distinguishes itself through its specialized training using the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models".

Key Capabilities

  • Enhanced Mathematical Reasoning: The primary focus of this model's fine-tuning is to improve its ability to handle complex mathematical problems and reasoning tasks.
  • Instruction Following: As a fine-tuned version of an instruct model, it is designed to follow user instructions effectively.
  • Contextual Understanding: With a substantial context length of 32768 tokens, it can process and generate responses based on extensive input.

Training Details

The model was fine-tuned using the TRL (Transformer Reinforcement Learning) library. The GRPO method, central to its training, aims to push the boundaries of mathematical reasoning in open-source language models. This approach suggests a strong performance in areas requiring logical deduction and numerical understanding.

Good For

  • Applications requiring robust mathematical problem-solving.
  • Tasks that benefit from advanced reasoning capabilities.
  • Scenarios where a smaller, yet capable, model with strong instruction-following is preferred for reasoning-intensive queries.