Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1
Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1 is an 8 billion parameter QLoRA fine-tune of Qwen3-8B, developed as a research artifact to study training-corpus overlap in medical language model benchmarks. This model is specifically optimized for medical question answering, demonstrating a +5.2 percentage point gain over its parent checkpoint on a held-out medical benchmark under overlap-audited evaluation. Its primary purpose is to enable reproducibility for a methodological paper on benchmark evaluation, highlighting how training-corpus exposure can inflate reported performance in medical LLMs.
Loading preview...
Model Overview
This model, Diagnostic-Reasoning-Q3X1, is an 8 billion parameter QLoRA fine-tune of Qwen3-8B, developed by Clinical-Reasoning-Hub. It serves as a research artifact for a study investigating the impact of training-corpus overlap on medical language model benchmarks. Initially, the model showed a significant gain, but subsequent overlap-audited evaluation revealed a more modest, yet statistically significant, +5.2 percentage point gain over the parent Qwen3-8B checkpoint on the MedXpertQA benchmark (exact-stem-unflagged subset).
Key Findings & Capabilities
- Methodological Contribution: The accompanying paper demonstrates that standard benchmark evaluation can conflate reasoning with exposure to test items, especially when fine-tuning corpora include public medical QA datasets.
- Overlap Auditing: The model's evaluation protocol includes a rigorous three-layer pipeline to measure benchmark overlap against its 97,041-item fine-tuning corpus, revealing significant overlap (e.g., 100% for PubMedQA, 73.4% for MedQA).
- Performance on Unflagged Data: On the MedXpertQA benchmark, which was not among the named training sources, the model achieved 20.8% accuracy on the exact-stem-unflagged subset, compared to 15.6% for the parent checkpoint.
- Training Methodology: Fine-tuned using QLoRA with a structured eight-component clinical reasoning template and a quality-tier-weighted curriculum, targeting
q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_projmodules.
Limitations & Intended Use
- Research Artifact: This model is explicitly a research artifact for reproducibility and has not been clinically validated. It must not be used for clinical decision-making.
- Evaluation Scope: Evaluated exclusively on English-language multiple-choice benchmarks, primarily MedXpertQA, with limitations regarding multilingual or cross-demographic assessment.
- Reproducibility: While accuracy results and statistical comparisons are reproducible, the training corpus itself is not redistributable, limiting independent regeneration of the overlap analysis.
This model is valuable for researchers interested in the integrity of LLM evaluation in specialized domains, particularly medicine, and for understanding the effects of data contamination.