luckeciano/Qwen-2.5-7B-DrGRPO-Adam-HessianMaskToken-1e-4-v3_6137

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 15, 2025Architecture:Transformer Featherless Exclusive Cold

The luckeciano/Qwen-2.5-7B-DrGRPO-Adam-HessianMaskToken-1e-4-v3_6137 is a 7.6 billion parameter language model, fine-tuned from Qwen/Qwen2.5-Math-7B, with a context length of 32768 tokens. It was trained using the GRPO method, as introduced in the DeepSeekMath paper, on the DigitalLearningGmbH/MATH-lighteval dataset. This model is specifically optimized for advanced mathematical reasoning tasks, leveraging its specialized training to enhance performance in complex quantitative problem-solving.

Loading preview...

Overview

This model, luckeciano/Qwen-2.5-7B-DrGRPO-Adam-HessianMaskToken-1e-4-v3_6137, is a 7.6 billion parameter language model built upon the Qwen/Qwen2.5-Math-7B base. It has been fine-tuned using the TRL framework, specifically incorporating the GRPO (Generalized Reinforcement Learning for Policy Optimization) method. This training approach, detailed in the DeepSeekMath paper, aims to push the boundaries of mathematical reasoning capabilities in open language models.

Key Capabilities

  • Enhanced Mathematical Reasoning: Optimized for solving complex mathematical problems, leveraging the GRPO training method.
  • Specialized Fine-tuning: Trained on the DigitalLearningGmbH/MATH-lighteval dataset, focusing on mathematical tasks.
  • Large Context Window: Supports a context length of 32768 tokens, allowing for processing longer and more intricate problem descriptions.

Good For

  • Mathematical Problem Solving: Ideal for applications requiring advanced quantitative analysis and step-by-step mathematical reasoning.
  • Research in LLM Optimization: Useful for researchers exploring the impact of GRPO and similar reinforcement learning techniques on model performance.
  • Educational Tools: Can be integrated into platforms designed to assist with or generate solutions for mathematical challenges.