sfewf/qwen3-4b-math-RL

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 1, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The sfewf/qwen3-4b-math-RL model is a Qwen3-4b variant specifically post-RL trained on a Math dataset, featuring a 'Max-Thinking mode' for enhanced reasoning. This model is optimized for mathematical problem-solving and complex reasoning tasks, demonstrating improved efficiency in both standard and max-thinking modes. It excels in accuracy on benchmarks like GSM8K, MATH-lighteval, BBH, and GPQA, making it suitable for applications requiring robust analytical capabilities.

Loading preview...

Model Overview

This model, sfewf/qwen3-4b-math-RL, is a specialized variant of the Qwen3-4b architecture that has undergone Reinforcement Learning (RL) training on a comprehensive Math dataset. A key feature is its Max-Thinking mode, inspired by DeepSeek V4, which encourages thorough, step-by-step deliberation for complex problems.

Key Capabilities

  • Enhanced Mathematical Reasoning: Demonstrates significant improvements in solving mathematical problems, as evidenced by benchmark scores.
  • Max-Thinking Mode: Provides a detailed, rigorous deliberation process, documenting intermediate steps and considered alternatives.
  • Efficient Reasoning: Observations indicate more efficient reasoning in both default and Max-Thinking modes.
  • Strong Benchmark Performance: Achieves high accuracy on various datasets:
    • GSM8K: 0.9172 (standard), 0.9327 (max-effort)
    • MATH-lighteval: 0.8019 (standard), 0.8505 (max-effort)
    • BBH: 0.7963 (standard), 0.8709 (max-effort)
    • GPQA: 0.2667 (standard), 0.3125 (max-effort)
  • Length-Penalty Training: Responses are concise while maintaining accuracy in default mode.

Good For

  • Mathematical Problem Solving: Ideal for applications requiring precise and detailed solutions to math-related queries.
  • Complex Reasoning Tasks: Benefits from the Max-Thinking mode for problems demanding extensive logical decomposition and verification.
  • Educational Tools: Can be used in systems that require step-by-step explanations for mathematical concepts.
  • Research and Development: Suitable for exploring advanced reasoning capabilities in LLMs.