gguk2on/qwen2.5-7B-rlar_g32_b384_math_vu2
TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026Architecture:Transformer Featherless Exclusive Cold
The gguk2on/qwen2.5-7B-rlar_g32_b384_math_vu2 model is a 7.6 billion parameter language model fine-tuned from Qwen/Qwen2.5-7B. It was trained using the GRPO method, specifically optimized for mathematical reasoning tasks. This model is designed to enhance performance in complex mathematical problem-solving, building upon the base Qwen2.5-7B architecture.
Loading preview...
Model Overview
This model, gguk2on/qwen2.5-7B-rlar_g32_b384_math_vu2, is a specialized fine-tuned version of the Qwen/Qwen2.5-7B base model, featuring 7.6 billion parameters and a 32k context length. It was developed by gguk2on and trained using the TRL framework.
Key Capabilities
- Enhanced Mathematical Reasoning: The model's primary differentiator is its fine-tuning with the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the DeepSeekMath research paper. This training approach specifically targets and improves the model's ability to handle complex mathematical problems and reasoning tasks.
- Qwen2.5-7B Foundation: Benefits from the robust capabilities of the Qwen2.5-7B architecture, providing a strong general language understanding base.
Training Details
- The training procedure leveraged the TRL library (version 0.16.0.dev0) and was tracked using Weights & Biases.
- The GRPO method, detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300), was central to its optimization for mathematical tasks.
Good For
- Applications requiring advanced mathematical problem-solving.
- Research and development in improving LLM performance on quantitative reasoning benchmarks.
- Use cases where a strong foundation model with specialized mathematical capabilities is beneficial.