Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Feb 21, 2026License:cc-by-nc-4.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1 is an 8 billion parameter QLoRA fine-tune of Qwen3-8B, developed for medical question answering. This model serves as a research artifact to study training-corpus overlap in medical language model benchmarks, demonstrating a +5.2 percentage point gain on a held-out benchmark subset after accounting for overlap. It is specifically designed to reproduce the results of an accompanying paper on methodological evaluation challenges in medical LLMs.

Loading preview...

What is Diagnostic-Reasoning-Q3X1?

This model is an 8 billion parameter QLoRA fine-tune of Qwen3-8B, specifically developed by Clinical-Reasoning-Hub for medical question answering. It functions primarily as a research artifact to support a study on the impact of training-corpus overlap in medical language model benchmarks. The model's release allows for independent reproduction of the paper's findings, which highlight how standard benchmark evaluation can conflate reasoning with exposure to test items when fine-tuning corpora include public medical QA datasets.

Key Characteristics & Performance

  • Focus on Reproducibility: The model's main purpose is to enable the reproduction of results from an accompanying paper on methodological evaluation.
  • Overlap-Audited Evaluation: Initial evaluations showing a +24.2-point gain were superseded due to significant training-corpus overlap. Corrected figures show a +5.2 percentage point gain over the parent Qwen3-8B checkpoint on the held-out, exact-stem-unflagged subset of the MedXpertQA benchmark.
  • Methodological Contribution: The model demonstrates that benchmark evaluation cannot reliably distinguish reasoning from exposure to test items when there is training-corpus overlap, with some benchmarks showing up to 100% overlap.
  • Training Details: Fine-tuned using QLoRA with r=128, α=256 on 349 million trainable parameters (4.09% of total) and a 97,041-example corpus, employing a structured eight-component clinical reasoning template.

Limitations and Intended Use

  • Research Artifact Only: This model has not been clinically validated and must not be used for clinical decision-making.
  • Evaluation Protocol Sensitivity: Performance can vary significantly based on inference engine builds and request-stream protocols, emphasizing the need for precise environment specification for reproducibility.
  • Limited Generalizability: Evaluated exclusively on English-language multiple-choice benchmarks, with training data predominantly from North American and UK clinical examination material. Conclusions primarily rest on the MedXpertQA benchmark, which was not among the named training sources.

Should I use this for my use case?

  • Good for:
    • Researchers interested in reproducing the findings of the associated paper on training-corpus overlap in medical LLM benchmarks.
    • Studying the methodological challenges of evaluating medical language models.
    • Understanding the impact of benchmark contamination on reported performance gains.
  • Not good for:
    • Clinical decision-making or any real-world medical applications.
    • General-purpose medical question answering without understanding its specific research context and limitations.
    • Applications requiring multilingual support or evaluation beyond English multiple-choice formats.