cjiao/goldengoose-divsweep_goose_n512_indorc_tau1.00-7grp

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 15, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweep_goose_n512_indorc_tau1.00-7grp model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, this model was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. With a context length of 32768 tokens, it is optimized for tasks requiring robust logical and mathematical problem-solving.

Loading preview...

Model Overview

The cjiao/goldengoose-divsweep_goose_n512_indorc_tau1.00-7grp is a 1.5 billion parameter instruction-tuned language model, building upon the Qwen/Qwen2.5-1.5B-Instruct architecture. It was fine-tuned by cjiao using the TRL framework.

Key Capabilities

  • Enhanced Mathematical Reasoning: This model was specifically trained with the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This training approach aims to significantly improve its ability to handle complex mathematical problems and logical deductions.
  • Instruction Following: As an instruction-tuned model, it is designed to follow user prompts and generate relevant responses effectively.
  • Large Context Window: It supports a substantial context length of 32768 tokens, allowing for processing and generating longer, more complex inputs and outputs.

Good For

  • Mathematical Problem Solving: Ideal for applications requiring strong mathematical reasoning, calculations, and logical inference.
  • Complex Instruction Following: Suitable for tasks where detailed and multi-step instructions need to be accurately interpreted and executed.
  • Research and Development: Provides a foundation for further experimentation and fine-tuning on tasks that benefit from advanced reasoning capabilities, particularly in the mathematical domain.