ExploreXploitQ/DeepMath-103K-Level6-Qwen3-4B-Base-GRPO

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ExploreXploitQ/DeepMath-103K-Level6-Qwen3-4B-Base-GRPO is a 4 billion parameter language model based on Qwen3-4B-Base, specifically post-trained for advanced mathematical reasoning. It utilizes Group Relative Policy Optimization (GRPO) on the Level 6 subset of the DeepMath-103K dataset. This model is optimized for solving complex math problems, demonstrated by its performance on AIME and AMC validation sets. With a context length of 32,768 tokens, it is designed for applications requiring deep mathematical problem-solving capabilities.

Loading preview...

Model Overview

ExploreXploitQ/DeepMath-103K-Level6-Qwen3-4B-Base-GRPO is a specialized 4-billion parameter language model derived from Qwen/Qwen3-4B-Base. Its primary distinction lies in its post-training methodology: it was fine-tuned using Group Relative Policy Optimization (GRPO) on the challenging Level 6 subset of the DeepMath-103K dataset.

Key Capabilities and Training

This model is engineered for advanced mathematical reasoning. The training involved 57,046 examples from DeepMath Level 6, utilizing the verl framework and full-parameter actor updates with rule-based math outcome rewards. Key training parameters include a context length of 32,768 tokens, 8 responses per prompt during rollout, and a constant learning rate of 1e-6 over one epoch.

Performance Highlights

Validation was conducted on competitive math datasets, yielding the following average accuracies at 16 responses:

  • AIME 2025: 21.25%
  • AMC 2022–2023: 60.77%
  • AIME 2024: 23.54%

Use Cases

This model is particularly well-suited for applications requiring robust mathematical problem-solving, especially those involving complex, multi-step reasoning found in advanced mathematics competitions or educational tools. Its specialized training makes it a strong candidate for tasks where precise mathematical output and step-by-step reasoning are critical.