TanitAI/Tanit-Med-8B-DPO

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Feb 25, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

TanitAI/Tanit-Med-8B-DPO is an 8.2 billion parameter medical reasoning model, fine-tuned from Qwen3-8B by TanitAI. It is specifically optimized for clinical multiple-choice reasoning and medical question answering, demonstrating strong performance on the MedAgentsBench standard splits. This model excels at USMLE-style vignettes and medical exam reasoning, featuring a unique DPO phase for calibrated confidence.

Loading preview...

Tanit-Med-8B-DPO: An 8B Medical Reasoning Model

Tanit-Med-8B-DPO is an 8.2 billion parameter model, fine-tuned from Qwen3-8B, designed for clinical multiple-choice reasoning and medical question answering. Developed by TanitAI, it underwent a four-stage training process including broad medical SFT, reasoning SFT, DPO, and a chain-of-thought polish. This model is notable for its performance on the MedAgentsBench, achieving a 50.7% macro average on standard splits, outperforming other open 8B medical models.

Key Capabilities & Features

  • Superior Medical Reasoning: Achieves leading scores among open 8B models on MedAgentsBench standard splits, particularly strong in MedQA, PubMedQA, and AfriMedQA.
  • DPO for Calibrated Confidence: The DPO phase (Phase 3) optimizes for preference alignment, resulting in fewer confidently wrong answers, trading slight raw accuracy for better-calibrated confidence.
  • Chain-of-Thought (CoT) Polish: Phase 4 recovers accuracy and ensures reasoning blocks (<think> blocks) are almost always present (≈99% of responses) and concise.
  • Optimized for Specific Formats: Best at multiple-choice clinical QA, USMLE-style vignettes, and medical exam reasoning.
  • North African Context Awareness: Shows significantly better performance on AfriMedQA compared to baselines, indicating improved grounding in African clinical practice.

Good for

  • Medical NLP research and benchmark development.
  • Medical education tooling and retrieval-augmented clinical QA prototypes.
  • As a base for further fine-tuning in medical domains.

Limitations

  • Not intended for clinical decision support, diagnosis, treatment, or any patient-facing deployment.
  • Struggles with MedXpertQA, indicating limitations at this scale for complex, multi-option questions.
  • English-only evaluation; long-form clinical advice is untested.
  • Inherits biases from Western and exam-derived training corpora and may hallucinate information.