logan7000/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-granite2b-math345-groupA-qwen25-end
This model is a 3.1 billion parameter language model fine-tuned from Qwen/Qwen2.5-3B, utilizing the GRPO method for enhanced mathematical reasoning. It is designed to improve performance on complex mathematical tasks, building upon the base Qwen2.5-3B architecture. With a context length of 32768 tokens, it is suitable for applications requiring robust mathematical problem-solving capabilities.
Loading preview...
Model Overview
This model is a fine-tuned variant of the Qwen2.5-3B architecture, specifically developed to enhance its mathematical reasoning abilities. It leverages the GRPO (Grouped Reinforcement Learning with Policy Optimization) training method, as introduced in the DeepSeekMath paper, to achieve improved performance in mathematical contexts.
Key Capabilities
- Enhanced Mathematical Reasoning: Trained with the GRPO method, which is designed to push the limits of mathematical problem-solving in language models.
- Base Model: Built upon the robust Qwen2.5-3B foundation, providing a strong general language understanding.
- Training Framework: Utilizes the TRL (Transformers Reinforcement Learning) library for its fine-tuning process.
Training Details
The model's training procedure involved the GRPO method, detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models." The training was conducted using TRL version 1.2.0.dev0, Transformers 4.57.6, Pytorch 2.10.0+cu128, Datasets 5.0.1, and Tokenizers 0.22.2.
Use Cases
This model is particularly well-suited for applications requiring advanced mathematical understanding and problem-solving, benefiting from its specialized GRPO-based fine-tuning.