narcolepticchicken/occ-grpo-v2-occ
narcolepticchicken/occ-grpo-v2-occ is a 3.1 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-3B-Instruct. This model was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is optimized for tasks requiring improved logical and mathematical problem-solving, building upon the base Qwen2.5 architecture.
Loading preview...
Model Overview
narcolepticchicken/occ-grpo-v2-occ is a 3.1 billion parameter language model, fine-tuned from the Qwen/Qwen2.5-3B-Instruct base model. It leverages the GRPO (Gradient-based Reward Policy Optimization) training method, as introduced in the DeepSeekMath paper, to enhance its reasoning abilities.
Key Capabilities
- Enhanced Reasoning: Specifically trained with GRPO, a method known for improving mathematical reasoning in language models.
- Instruction Following: Builds upon the instruction-tuned capabilities of the Qwen2.5-3B-Instruct base model.
- Context Length: Supports a context length of 32768 tokens, allowing for processing longer inputs.
Training Details
The model was fine-tuned using the TRL (Transformers Reinforcement Learning) framework. The GRPO method, which is central to its training, aims to push the limits of mathematical reasoning in open language models, as detailed in the DeepSeekMath paper.
Good For
- Applications requiring improved logical and mathematical problem-solving.
- Tasks where enhanced reasoning capabilities are beneficial, especially in a smaller parameter count model.
- Developers looking for a Qwen2.5-based model with specialized reasoning fine-tuning.