luckeciano/Qwen-2.5-7B-GRPO-Base-KL-0.01-v2_6050
TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 15, 2025Architecture:Transformer Featherless Exclusive Cold
The luckeciano/Qwen-2.5-7B-GRPO-Base-KL-0.01-v2_6050 is a 7.6 billion parameter language model, fine-tuned from Qwen/Qwen2.5-Math-7B, with a context length of 32768 tokens. It was trained using the GRPO method on the DigitalLearningGmbH/MATH-lighteval dataset, specializing it in mathematical reasoning tasks. This model is optimized for advanced mathematical problem-solving, leveraging techniques from DeepSeekMath to enhance its reasoning capabilities.
Loading preview...
Overview
This model, luckeciano/Qwen-2.5-7B-GRPO-Base-KL-0.01-v2_6050, is a 7.6 billion parameter language model derived from the Qwen/Qwen2.5-Math-7B base model. It has been specifically fine-tuned using the TRL library on the DigitalLearningGmbH/MATH-lighteval dataset.
Key Capabilities
- Enhanced Mathematical Reasoning: The model's training incorporates the GRPO (Generalized Reinforcement Learning with Policy Optimization) method, as introduced in the DeepSeekMath research, to significantly improve its ability to handle complex mathematical problems.
- Specialized Fine-tuning: By focusing on a dedicated mathematical dataset, this model is tailored for tasks requiring precise numerical and logical deduction.
Good For
- Mathematical Problem Solving: Ideal for applications requiring robust mathematical reasoning, from algebra to advanced calculus.
- Research and Development: Suitable for researchers exploring advanced reinforcement learning techniques in language models, particularly those interested in GRPO.
- Educational Tools: Can be integrated into systems designed to assist with or generate solutions for mathematical challenges.