pb09204048/Qwen3-4B-Math-RL-Specialist

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 17, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The pb09204048/Qwen3-4B-Math-RL-Specialist is a 4 billion parameter BF16 math specialist model, fine-tuned from Qwen/Qwen3-4B using a rank-16 LoRA. It is specifically optimized for mathematical reasoning tasks, demonstrating a significant performance increase on the AIME24 benchmark. This model is designed to serve as a domain teacher for Multi-Teacher On-Policy Distillation (MOPD) experiments, focusing on mathematical problem-solving.

Loading preview...

Overview

This model, pb09204048/Qwen3-4B-Math-RL-Specialist, is a 4 billion parameter BF16 math specialist derived from the Qwen3-4B base model. It was fine-tuned using a rank-16 LoRA and merged into full Hugging Face weights. The primary purpose of this specialist is to act as a domain teacher for Multi-Teacher On-Policy Distillation (MOPD) experiments, specifically in the mathematical domain.

Key Capabilities & Training

  • Mathematical Specialization: Trained extensively on the zhuzilin/dapo-math-17k dataset, comprising 15,981 unique math prompts.
  • Performance: Achieved a 50.00% pass@1 score on the AIME24 benchmark, a substantial improvement over the Qwen3-4B base model's 22.50%.
  • Training Method: Utilizes independent domain Reinforcement Learning (RL) with GRPO-style group-centered advantages and a PPO clipped policy objective. The training involved 500 optimizer updates.
  • Context Length: Supports a maximum response length of 16,384 tokens and a 24,576-token context during evaluation.

Use Cases

This model is particularly suited for:

  • MOPD Experiments: Designed to be used as a math teacher model within a MOPD framework, providing token-level supervision for student models.
  • Mathematical Reasoning: Excels at solving complex mathematical problems, as evidenced by its AIME24 performance.
  • Research: Ideal for researchers exploring domain-specific RL fine-tuning and multi-teacher distillation techniques.