cmcheng/DeepMath-GRPO_Qwen2.5-0.5B-Instruct

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 10, 2026Architecture:Transformer Featherless Exclusive Cold

cmcheng/DeepMath-GRPO_Qwen2.5-0.5B-Instruct is a Qwen2.5-0.5B-Instruct model fine-tuned by cmcheng using the GRPO (Generalized Reinforcement Learning with Policy Optimization) method. This model is specifically optimized for mathematical reasoning tasks, leveraging the DeepMath-103K dataset. It aims to enhance performance on complex mathematical problems through reinforcement learning techniques.

Loading preview...

Model Overview

cmcheng/DeepMath-GRPO_Qwen2.5-0.5B-Instruct is a specialized language model based on the Qwen2.5-0.5B-Instruct architecture. It has been fine-tuned using the GRPO (Generalized Reinforcement Learning with Policy Optimization) method, a reinforcement learning approach, to significantly improve its capabilities in mathematical reasoning.

Key Capabilities

  • Mathematical Reasoning: Optimized for solving complex mathematical problems, leveraging a dedicated dataset.
  • GRPO Fine-tuning: Utilizes the GRPO algorithm with specific parameters for learning, including a learning rate of 1e-6, KL divergence control (beta=0.001), and clipping parameters (epsilon=0.2, epsilon_high=0.28).
  • Training Data: Trained on the zwhe99/DeepMath-103K dataset, comprising 97,870 training samples.
  • Efficient Training: Trained using DeepSpeed with bf16 mixed-precision on NVIDIA 4080 GPUs, incorporating gradient accumulation and checkpointing for memory optimization.

Good For

This model is particularly well-suited for applications requiring strong mathematical problem-solving abilities. Its GRPO fine-tuning makes it a candidate for tasks where traditional instruction-tuned models might struggle with the nuances of mathematical logic and reasoning.