zou-lab/BioMed-R1-32B

TEXT GENERATIONConcurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 25, 2025License:llama3.1Architecture:Transformer0.0K Featherless Exclusive Cold

BioMed-R1-32B is a 32.8 billion parameter medical large language model developed by zou-lab, fine-tuned using supervised learning and reinforcement learning on reasoning-heavy and adversarial examples. It is designed to improve medical reasoning capabilities and encourage self-correction and backtracking in diagnostic scenarios. The model achieves strong overall and adversarial performance among similarly sized biomedical LLMs, particularly in disentangling reasoning from knowledge-based questions. It is built upon the Qwen2.5-32B-Instruct architecture and supports a 32768 token context length.

Loading preview...

BioMed-R1-32B: Enhanced Medical Reasoning LLM

BioMed-R1-32B is a 32.8 billion parameter medical large language model from zou-lab, specifically developed to address the challenges of medical reasoning in LLMs. Traditional medical benchmarks often mix reasoning-heavy and knowledge-heavy questions, making it difficult to accurately evaluate a model's true reasoning abilities. This model leverages a novel approach to disentangle these question types, revealing that a significant portion of medical questions require complex reasoning rather than just factual recall.

Key Capabilities and Differentiators

  • Disentangled Reasoning Evaluation: Utilizes a PubMedBERT-based classifier to differentiate between reasoning-heavy and knowledge-heavy questions across 11 biomedical QA benchmarks, providing a more accurate assessment of reasoning performance.
  • Robustness to Adversarial Inputs: Trained with supervised fine-tuning and reinforcement learning on adversarial examples, BioMed-R1-32B is designed to encourage self-correction and backtracking, making it more resilient to incorrect prefilled answers compared to other biomedical models.
  • Improved Medical Reasoning: Achieves strong overall and adversarial performance among similarly sized biomedical LLMs by focusing on training strategies that promote reasoning under uncertainty.

Use Cases and Strengths

  • Medical Question Answering: Excels in scenarios requiring complex medical reasoning, moving beyond simple factual recall.
  • Diagnostic Support: Offers enhanced robustness in challenging diagnostic contexts where models might encounter misleading information.
  • Research and Development: Provides a foundation for further research into improving medical LLM reasoning, particularly by incorporating reasoning-rich data sources like clinical case reports.