Jongbin-kr/qwen2.5-coder-7b-verireason-grpo-official-full-ft
Jongbin-kr/qwen2.5-coder-7b-verireason-grpo-official-full-ft is a 7.6 billion parameter language model fine-tuned from Jongbin-kr/qwen2.5-coder-7b-verireason_sft-reasoning_official-full-ft. This model was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is optimized for tasks requiring advanced reasoning, particularly in mathematical contexts, building upon its coder base.
Loading preview...
Model Overview
This model, Jongbin-kr/qwen2.5-coder-7b-verireason-grpo-official-full-ft, is a 7.6 billion parameter language model. It is a fine-tuned iteration of the Jongbin-kr/qwen2.5-coder-7b-verireason_sft-reasoning_official-full-ft base model.
Key Differentiator
The primary distinction of this model lies in its training methodology. It was fine-tuned using GRPO (Gradient Regularized Policy Optimization), a method introduced in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This indicates a specific optimization for enhancing mathematical reasoning capabilities.
Training Details
- Methodology: Trained with GRPO using the TRL framework (version 1.6.0).
- Framework Versions: Utilizes Transformers 5.7.0, Pytorch 2.10.0+cu128, Datasets 5.0.0, and Tokenizers 0.22.2.
Use Cases
Given its GRPO training for mathematical reasoning, this model is particularly suited for:
- Tasks requiring complex mathematical problem-solving.
- Applications where robust logical and reasoning abilities are crucial.
- Scenarios benefiting from enhanced numerical and symbolic manipulation.