q1716523669/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-phi4mini-math345-groupC-phi4mini-end
The q1716523669/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-phi4mini-math345-groupC-phi4mini-end model is a 3.8 billion parameter language model fine-tuned from Microsoft's Phi-4-mini-instruct. It was trained using the GRPO (Gradient-based Reward Policy Optimization) method, which is designed to enhance mathematical reasoning capabilities. This model is particularly suited for tasks requiring advanced mathematical problem-solving and logical deduction, building upon the base Phi-4-mini-instruct architecture.
Loading preview...
Model Overview
This model, developed by q1716523669, is a fine-tuned version of the microsoft/Phi-4-mini-instruct base model, featuring approximately 3.8 billion parameters and a 32768 token context length. It leverages the TRL (Transformers Reinforcement Learning) framework for its training process.
Key Differentiator: GRPO Training
A significant aspect of this model is its training methodology, which incorporates GRPO (Gradient-based Reward Policy Optimization). This technique, introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models," is specifically designed to improve a model's mathematical reasoning abilities. This suggests the model is optimized for complex numerical and logical tasks.
Potential Use Cases
- Mathematical Problem Solving: Ideal for applications requiring robust mathematical reasoning.
- Logical Deduction: Can be applied to tasks that benefit from enhanced logical processing.
- Instruction Following: Benefits from its
Phi-4-mini-instructbase, making it suitable for instruction-tuned applications.
Technical Details
The model was trained using TRL version 1.2.0.dev0, with Transformers 4.57.6, Pytorch 2.10.0+cu128, Datasets 5.0.1, and Tokenizers 0.22.2. The training process can be visualized via Weights & Biases.