cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_seed100_random_n25

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 28, 2026Architecture:Transformer Featherless Exclusive Cold

The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_seed100_random_n25 model is a 1.5 billion parameter language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct with a 32K context length. It was trained using the GRPO method, as introduced in the DeepSeekMath paper, which focuses on enhancing mathematical reasoning. This model is optimized for tasks requiring improved reasoning capabilities, particularly in mathematical contexts, building upon the Qwen2.5 architecture.

Loading preview...

Overview

This model, cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_seed100_random_n25, is a 1.5 billion parameter language model derived from the Qwen/Qwen2.5-1.5B-Instruct base. It features a substantial context length of 32,768 tokens, making it suitable for processing longer inputs.

Key Capabilities & Training

The model was fine-tuned using the TRL framework and incorporates the GRPO (Gradient-based Reward Policy Optimization) method. GRPO is a technique highlighted in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This training approach suggests an emphasis on improving the model's reasoning abilities, particularly in mathematical domains.

Use Cases

Given its foundation in Qwen2.5-1.5B-Instruct and its GRPO-based training, this model is likely well-suited for applications requiring enhanced reasoning, especially in areas that benefit from structured problem-solving or mathematical understanding. Developers can integrate it using the Hugging Face transformers library for text generation tasks.