cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_indorc_tau0.10_n25

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 21, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_indorc_tau0.10_n25 is a 1.5 billion parameter language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, this model utilizes the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is optimized for tasks requiring advanced mathematical understanding and problem-solving, making it suitable for applications in scientific computing and quantitative analysis.

Loading preview...

Overview

This model, goldengoose-divsweepv2_lowdiv_goose_n128_indorc_tau0.10_n25, is a 1.5 billion parameter language model fine-tuned by cjiao. It is based on the Qwen/Qwen2.5-1.5B-Instruct architecture and was trained using the TRL framework.

Key Differentiator

The primary distinction of this model lies in its training methodology. It incorporates the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This technique is specifically designed to improve a model's proficiency in mathematical reasoning tasks.

Training Details

  • Base Model: Qwen/Qwen2.5-1.5B-Instruct
  • Training Framework: TRL (Transformer Reinforcement Learning)
  • Optimization Method: GRPO, focusing on enhancing mathematical reasoning.

Potential Use Cases

  • Mathematical Problem Solving: Ideal for applications requiring robust mathematical reasoning.
  • Scientific Computing: Can be leveraged in fields that demand precise quantitative analysis.
  • Educational Tools: Development of AI tutors or assistants for math and science.

This model offers a specialized approach to language understanding, particularly in domains where mathematical accuracy and reasoning are paramount, setting it apart from general-purpose instruction-tuned models.