TanitAI/Tanit-Med-8B-DPO
TanitAI/Tanit-Med-8B-DPO is an 8.2 billion parameter medical reasoning model, fine-tuned from Qwen3-8B by TanitAI. It is specifically optimized for clinical multiple-choice reasoning and medical question answering, demonstrating strong performance on the MedAgentsBench standard splits. This model excels at USMLE-style vignettes and medical exam reasoning, featuring a unique DPO phase for calibrated confidence.
Loading preview...
Tanit-Med-8B-DPO: An 8B Medical Reasoning Model
Tanit-Med-8B-DPO is an 8.2 billion parameter model, fine-tuned from Qwen3-8B, designed for clinical multiple-choice reasoning and medical question answering. Developed by TanitAI, it underwent a four-stage training process including broad medical SFT, reasoning SFT, DPO, and a chain-of-thought polish. This model is notable for its performance on the MedAgentsBench, achieving a 50.7% macro average on standard splits, outperforming other open 8B medical models.
Key Capabilities & Features
- Superior Medical Reasoning: Achieves leading scores among open 8B models on MedAgentsBench standard splits, particularly strong in MedQA, PubMedQA, and AfriMedQA.
- DPO for Calibrated Confidence: The DPO phase (Phase 3) optimizes for preference alignment, resulting in fewer confidently wrong answers, trading slight raw accuracy for better-calibrated confidence.
- Chain-of-Thought (CoT) Polish: Phase 4 recovers accuracy and ensures reasoning blocks (
<think>blocks) are almost always present (≈99% of responses) and concise. - Optimized for Specific Formats: Best at multiple-choice clinical QA, USMLE-style vignettes, and medical exam reasoning.
- North African Context Awareness: Shows significantly better performance on AfriMedQA compared to baselines, indicating improved grounding in African clinical practice.
Good for
- Medical NLP research and benchmark development.
- Medical education tooling and retrieval-augmented clinical QA prototypes.
- As a base for further fine-tuning in medical domains.
Limitations
- Not intended for clinical decision support, diagnosis, treatment, or any patient-facing deployment.
- Struggles with MedXpertQA, indicating limitations at this scale for complex, multi-option questions.
- English-only evaluation; long-form clinical advice is untested.
- Inherits biases from Western and exam-derived training corpora and may hallucinate information.