gguk2on/qwen2.5-7B-rlar_g32_b384_math_vu2

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026Architecture:Transformer Featherless Exclusive Cold

The gguk2on/qwen2.5-7B-rlar_g32_b384_math_vu2 model is a 7.6 billion parameter language model fine-tuned from Qwen/Qwen2.5-7B. It was trained using the GRPO method, specifically optimized for mathematical reasoning tasks. This model is designed to enhance performance in complex mathematical problem-solving, building upon the base Qwen2.5-7B architecture.

Loading preview...

Model Overview

This model, gguk2on/qwen2.5-7B-rlar_g32_b384_math_vu2, is a specialized fine-tuned version of the Qwen/Qwen2.5-7B base model, featuring 7.6 billion parameters and a 32k context length. It was developed by gguk2on and trained using the TRL framework.

Key Capabilities

  • Enhanced Mathematical Reasoning: The model's primary differentiator is its fine-tuning with the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath research paper. This training approach specifically targets and improves the model's ability to handle complex mathematical problems and reasoning tasks.
  • Qwen2.5-7B Foundation: Benefits from the robust capabilities of the Qwen2.5-7B architecture, providing a strong general language understanding base.

Training Details

  • The training procedure leveraged the TRL library (version 0.16.0.dev0) and was tracked using Weights & Biases.
  • The GRPO method, detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300), was central to its optimization for mathematical tasks.

Good For

  • Applications requiring advanced mathematical problem-solving.
  • Research and development in improving LLM performance on quantitative reasoning benchmarks.
  • Use cases where a strong foundation model with specialized mathematical capabilities is beneficial.