cjiao/goldengoose-divsweep_goose_n128_indorc_tau0.10_seed100-25grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 26, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweep_goose_n128_indorc_tau0.10_seed100-25grp model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is particularly suited for tasks requiring improved logical and mathematical problem-solving, leveraging its specialized training approach.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweep_goose_n128_indorc_tau0.10_seed100-25grp is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. It utilizes a context length of 32768 tokens.

Key Capabilities

  • Enhanced Mathematical Reasoning: This model was specifically trained using the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath paper. This training approach aims to significantly improve the model's ability to handle mathematical and logical reasoning tasks.
  • Instruction Following: As a fine-tuned instruction model, it is designed to follow user prompts effectively, generating relevant and coherent responses.
  • TRL Framework: The fine-tuning process was conducted using the TRL (Transformer Reinforcement Learning) library, indicating a reinforcement learning approach to align the model with desired behaviors.

Training Details

The model's training procedure involved the GRPO method, which is detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models". This suggests a focus on robust and accurate mathematical problem-solving.

Good For

  • Applications requiring strong mathematical reasoning.
  • Tasks where logical problem-solving is critical.
  • Instruction-following scenarios benefiting from specialized mathematical training.