hkr04/qwen3-4b-grpo-dapo17k

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 13, 2026Architecture:Transformer Featherless Exclusive Cold

hkr04/qwen3-4b-grpo-dapo17k is a 4 billion parameter Qwen3-based language model. It has been fine-tuned using the GRPO method on the DAPO-Math-17k dataset. This model is specifically optimized for mathematical reasoning and problem-solving tasks, leveraging its training on a specialized math dataset.

Loading preview...

Model Overview

hkr04/qwen3-4b-grpo-dapo17k is a 4 billion parameter language model built upon the Qwen3 architecture. Its primary distinction lies in its specialized training regimen: it has been fine-tuned using the GRPO (Grouped Reinforcement Learning from Human Feedback with Policy Optimization) method. This process was applied to the DAPO-Math-17k dataset, indicating a strong focus on enhancing mathematical reasoning and problem-solving capabilities.

Key Training Details

  • Training Method: GRPO (Grouped Reinforcement Learning from Human Feedback with Policy Optimization)
  • Dataset: DAPO-Math-17k, a specialized dataset for mathematical tasks.
  • Batch Size: 32
  • Group Size: 8
  • Epochs: 1
  • Steps: 559
  • Maximum Response Length: 8192 tokens

Good For

  • Mathematical Reasoning: Excels in tasks requiring logical deduction and problem-solving within a mathematical context.
  • Specialized Math Applications: Suitable for applications where strong performance on mathematical datasets is crucial.
  • Research in RLHF for Math: Provides a base for further experimentation with GRPO and similar fine-tuning techniques on mathematical domains.