cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_indorc_seed100_tau0.10_n25
TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 28, 2026Architecture:Transformer Featherless Exclusive Cold
The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_indorc_seed100_tau0.10_n25 model is a 1.5 billion parameter language model fine-tuned by cjiao, based on Qwen/Qwen2.5-1.5B-Instruct. It was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring advanced mathematical problem-solving and logical deduction, building upon the robust Qwen2.5 architecture.
Loading preview...
Model Overview
The cjiao/goldengoose-divsweepv2_lowdiv_goose_n128_indorc_seed100_tau0.10_n25 is a 1.5 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-1.5B-Instruct base model. It leverages a 32K context length, making it suitable for processing longer inputs.
Key Capabilities
- Enhanced Mathematical Reasoning: This model was specifically trained using the GRPO (Generalized Reinforcement Learning for Policy Optimization) method, as introduced in the DeepSeekMath paper. This training approach aims to significantly improve its ability to handle complex mathematical problems and logical reasoning tasks.
- Instruction Following: As a fine-tuned version of an instruction-tuned model, it is designed to follow user instructions effectively.
- TRL Framework: The fine-tuning process utilized the TRL (Transformer Reinforcement Learning) library, indicating a focus on optimizing model behavior through reinforcement learning techniques.
Good For
- Mathematical Problem Solving: Ideal for applications requiring strong mathematical reasoning, such as solving equations, proofs, or quantitative analysis.
- Complex Logical Tasks: Suitable for scenarios where the model needs to deduce answers from intricate information or follow multi-step logical processes.
- Research and Development: Provides a foundation for further experimentation with GRPO-based fine-tuning or mathematical reasoning tasks on a Qwen2.5 base.