cjiao/goldengoose-divsweep_goose_n512_random_seed100-7grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 27, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweep_goose_n512_random_seed100-7grp model is a 1.5 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is particularly suited for tasks requiring improved logical and mathematical problem-solving, building upon its Qwen2.5 base.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweep_goose_n512_random_seed100-7grp is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. It leverages the TRL (Transformer Reinforcement Learning) framework for its training process.

Key Capabilities

  • Enhanced Mathematical Reasoning: This model was specifically trained using the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper. This training approach aims to significantly improve the model's ability to handle mathematical and logical reasoning tasks.
  • Instruction Following: As a fine-tuned version of an instruct model, it is designed to follow user instructions effectively.
  • Efficient Performance: With 1.5 billion parameters, it offers a balance between performance and computational efficiency, making it suitable for various applications where larger models might be overkill.

Training Details

The model's training procedure utilized GRPO, a method detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). The training environment included TRL 0.19.1, Transformers 4.57.6, Pytorch 2.5.1, Datasets 4.8.4, and Tokenizers 0.22.2.

Good For

  • Applications requiring strong mathematical problem-solving.
  • Tasks benefiting from improved logical reasoning.
  • Instruction-following scenarios where a 1.5B parameter model is sufficient.