cjiao/goldengoose-divsweepv2_lowdiv_goose_n512_random_n7

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 21, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n512_random_n7 model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the TRL library and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring improved reasoning, particularly in mathematical contexts, making it suitable for applications where robust logical processing is crucial.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n512_random_n7 is a 1.5 billion parameter language model, fine-tuned from the Qwen2.5-1.5B-Instruct architecture. This model leverages the TRL library for its training process.

Key Capabilities

  • Enhanced Reasoning: The model's training incorporates the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models." This suggests a focus on improving the model's ability to handle complex reasoning tasks.
  • Instruction Following: As a fine-tuned version of an instruction-tuned base model, it is designed to follow user instructions effectively.

Training Details

The model was trained using the GRPO method, which is detailed in the DeepSeekMath paper. This method aims to push the boundaries of mathematical reasoning in open language models. The training utilized specific versions of frameworks including TRL 0.19.1, Transformers 4.57.6, Pytorch 2.5.1, Datasets 4.8.4, and Tokenizers 0.22.2.

Good For

  • Applications requiring improved mathematical reasoning.
  • Tasks where robust instruction following is important.
  • Experimentation with models fine-tuned using advanced reinforcement learning techniques like GRPO.