AIArchiveInfo/Baichuan-M2-32B

TEXT GENERATIONPricing:Input $2.72 / Output $4.8Concurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 18, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Baichuan-M2-32B is a 32.8 billion parameter medical-enhanced reasoning model developed by Baichuan AI, built upon Qwen2.5-32B. It integrates a Large Verifier System and medical domain adaptation through Mid-Training to achieve breakthrough performance in real-world medical reasoning tasks. This model excels at clinical diagnostic thinking and patient interaction, making it suitable for medical education, health consultation, and clinical decision support.

Loading preview...

Baichuan-M2-32B: Medical-Enhanced Reasoning Model

Baichuan-M2-32B, developed by Baichuan AI, is a 32.8 billion parameter model specifically designed for medical reasoning tasks. Building on Qwen2.5-32B, it introduces a Large Verifier System and medical domain adaptation enhancement via Mid-Training to achieve high performance in medical scenarios while retaining strong general capabilities.

Key Innovations & Features

  • Large Verifier System: Incorporates a comprehensive medical verification framework with patient simulators and multi-dimensional verification (e.g., medical accuracy, response completeness).
  • Medical Domain Adaptation: Achieves efficient medical knowledge injection through Mid-Training and a multi-stage reinforcement learning strategy, balancing medical, general, and mathematical training data.
  • Doctor-Thinking Alignment: Trained on real clinical cases and patient simulators to develop clinical diagnostic thinking and robust patient interaction.
  • Efficient Deployment: Supports 4-bit quantization, enabling deployment on a single RTX4090, with improved token throughput in its MTP version for single-user scenarios.

Performance Highlights

Baichuan-M2-32B demonstrates leading performance among open-source medical models, achieving a HealthBench score of 60.1, which is noted as being closest to GPT-5's medical capabilities. It also shows strong general performance, outperforming Qwen3-32B (Thinking) on benchmarks like AIME24, Arena-Hard-v2.0, and CFBench.

Intended Use Cases

  • Medical Education: For learning and training purposes.
  • Health Consultation: Providing information and guidance.
  • Clinical Decision Support: Assisting medical professionals in their diagnostic and treatment processes.

Note: This model is intended for research and reference only and should not replace professional medical diagnosis or treatment. It is recommended for use under the guidance of medical professionals.