gguk2on/qwen2.5-7B-rlar_g8_b384_math_0.60.08
The gguk2on/qwen2.5-7B-rlar_g8_b384_math_0.60.08 is a 7.6 billion parameter language model fine-tuned from Qwen/Qwen2.5-7B. This model specializes in mathematical reasoning, having been trained with the GRPO method introduced in the DeepSeekMath paper. It is optimized for tasks requiring advanced mathematical problem-solving capabilities.
Loading preview...
Model Overview
This model, gguk2on/qwen2.5-7B-rlar_g8_b384_math_0.60.08, is a fine-tuned variant of the Qwen/Qwen2.5-7B base model, featuring 7.6 billion parameters and a 32K context length. It has been specifically enhanced for mathematical reasoning tasks.
Key Capabilities
- Mathematical Reasoning: The model's primary strength lies in its ability to handle complex mathematical problems, a result of its training with the GRPO method.
- Fine-tuned Performance: Built upon the robust Qwen2.5-7B architecture, it leverages advanced fine-tuning techniques using the TRL library.
Training Details
The model was trained using the GRPO (Gradient-based Reward Policy Optimization) method, as detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This specialized training approach aims to significantly improve its mathematical problem-solving skills. The training utilized TRL, Transformers, Pytorch, Datasets, and Tokenizers libraries, with specific versioning for reproducibility.