gguk2on/qwen2.5-7B-rlar_g8_b384_math_0.60.08

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 9, 2026Architecture:Transformer Featherless Exclusive Cold

The gguk2on/qwen2.5-7B-rlar_g8_b384_math_0.60.08 is a 7.6 billion parameter language model fine-tuned from Qwen/Qwen2.5-7B. This model specializes in mathematical reasoning, having been trained with the GRPO method introduced in the DeepSeekMath paper. It is optimized for tasks requiring advanced mathematical problem-solving capabilities.

Loading preview...

Model Overview

This model, gguk2on/qwen2.5-7B-rlar_g8_b384_math_0.60.08, is a fine-tuned variant of the Qwen/Qwen2.5-7B base model, featuring 7.6 billion parameters and a 32K context length. It has been specifically enhanced for mathematical reasoning tasks.

Key Capabilities

  • Mathematical Reasoning: The model's primary strength lies in its ability to handle complex mathematical problems, a result of its training with the GRPO method.
  • Fine-tuned Performance: Built upon the robust Qwen2.5-7B architecture, it leverages advanced fine-tuning techniques using the TRL library.

Training Details

The model was trained using the GRPO (Gradient-based Reward Policy Optimization) method, as detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This specialized training approach aims to significantly improve its mathematical problem-solving skills. The training utilized TRL, Transformers, Pytorch, Datasets, and Tokenizers libraries, with specific versioning for reproducibility.