cjiao/goldengoose-p3_goose_highdiv_n128_indoc_tau0.10-25grp
The cjiao/goldengoose-p3_goose_highdiv_n128_indoc_tau0.10-25grp model is a 1.5 billion parameter language model fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Developed by cjiao, this model was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. With a context length of 32768 tokens, it is optimized for tasks requiring advanced reasoning, particularly in mathematical domains, leveraging techniques from the DeepSeekMath research. This model is suitable for applications needing robust logical and mathematical problem-solving at a smaller scale.
Loading preview...
Model Overview
This model, goldengoose-p3_goose_highdiv_n128_indoc_tau0.10-25grp, is a 1.5 billion parameter language model built upon the Qwen/Qwen2.5-1.5B-Instruct architecture. It has been fine-tuned using the TRL framework.
Key Capabilities & Training
The primary differentiator of this model is its training methodology. It utilizes GRPO (Guided Reasoning Policy Optimization), a method detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This training approach suggests an emphasis on improving mathematical reasoning and problem-solving abilities.
Use Cases
Given its specialized training with GRPO, this model is particularly well-suited for:
- Mathematical problem-solving: Tasks that require logical deduction and numerical reasoning.
- Reasoning-intensive applications: Scenarios where robust analytical capabilities are crucial.
- Instruction following: Leveraging its base as an instruction-tuned model for diverse prompts.
This model offers a compact yet powerful option for developers seeking enhanced reasoning performance, especially in mathematical contexts, within a 1.5B parameter footprint.