cjiao/goldengoose-divsweepv2_lowdiv_goose_n512_grouporc_tau1.00_n7

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 21, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n512_grouporc_tau1.00_n7 model is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring robust reasoning, particularly in mathematical contexts, leveraging techniques from DeepSeekMath research. With a 32K context length, it aims to provide improved performance for complex problem-solving.

Loading preview...

Overview

This model, goldengoose-divsweepv2_lowdiv_goose_n512_grouporc_tau1.00_n7, is a 1.5 billion parameter instruction-tuned language model. It is a fine-tuned variant of the Qwen/Qwen2.5-1.5B-Instruct base model, developed by cjiao.

Key Capabilities & Training

  • Mathematical Reasoning Enhancement: The model was trained using the GRPO (Grouped Reinforcement Learning with Policy Optimization) method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models". This indicates a focus on improving its ability to handle complex mathematical problems and reasoning tasks.
  • Fine-tuning Framework: Training was conducted using the TRL (Transformer Reinforcement Learning) library, a framework from Hugging Face designed for fine-tuning large language models.
  • Context Length: It supports a substantial context window of 32,768 tokens, allowing for processing longer inputs and maintaining coherence over extended conversations or documents.

Intended Use Cases

This model is particularly well-suited for applications that require:

  • Mathematical problem-solving and reasoning.
  • Instruction following in complex scenarios, benefiting from its instruction-tuned base.
  • Tasks where a longer context window is advantageous for understanding and generating relevant responses.