logan7000/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-granite2b-math345-groupB-llama32-end
This model, named group_B, is a 3.2 billion parameter language model fine-tuned from meta-llama/Llama-3.2-3B-Instruct. It was trained using the TRL framework and incorporates the GRPO method, which is designed to enhance mathematical reasoning capabilities. This model is specifically optimized for tasks requiring advanced mathematical problem-solving and logical deduction.
Loading preview...
Overview
This model, designated as group_B, is a 3.2 billion parameter language model that has been fine-tuned from the meta-llama/Llama-3.2-3B-Instruct base model. The fine-tuning process utilized the TRL (Transformers Reinforcement Learning) framework.
Key Capabilities
- Enhanced Mathematical Reasoning: A core differentiator of this model is its training with the GRPO (Gradient-based Reinforcement Learning for Policy Optimization) method. This technique, introduced in the DeepSeekMath paper, is specifically aimed at improving the model's ability to handle complex mathematical reasoning tasks.
- Instruction Following: As it is fine-tuned from an instruct model, it is designed to follow user instructions effectively.
Training Details
The model's training procedure leveraged the GRPO method, as detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). The training environment included specific versions of key frameworks:
- TRL: 1.2.0.dev0
- Transformers: 4.57.6
- Pytorch: 2.10.0+cu128
- Datasets: 5.0.1
- Tokenizers: 0.22.2
Good For
- Applications requiring strong mathematical problem-solving.
- Tasks that benefit from advanced logical reasoning.
- Developers looking for a compact yet capable model for instruction-tuned mathematical and reasoning-based queries.