Salahuddin1234/omnidoctor-MedPsy

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 30, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

MedPsy-4B is a 4 billion parameter text-only causal language model developed by Tether AI Research, built on Qwen3-4B-Thinking-2507. It is specifically designed for medical and healthcare applications, excelling on medical benchmarks and outperforming models nearly 7x its size. This model is optimized for efficient edge deployment, offering significant token efficiency for real-time clinical decision support.

Loading preview...

MedPsy-4B: A Compact Medical LLM for Edge Devices

MedPsy-4B, developed by Tether AI Research, is a 4 billion parameter text-only causal language model built upon the Qwen3-4B-Thinking-2507 architecture. It has been post-trained using a multi-stage pipeline involving supervised fine-tuning and reinforcement learning on curated medical data, specifically optimized for medical and healthcare applications.

Key Capabilities & Differentiators

  • Exceptional Performance at Scale: Achieves an average score of 70.54 on closed-ended medical benchmarks, surpassing MedGemma-27B-text-it (69.95) despite being nearly seven times smaller.
  • Strong Clinical Relevance: Scores 74.00 on HealthBench and 58.00 on HealthBench Hard, outperforming MedGemma-27B by significant margins.
  • High Token Efficiency: Demonstrates a 3.2x reduction in average response length compared to its base model, leading to faster inference and lower compute costs, crucial for edge deployment.
  • Privacy-First Design: Supports fully on-device inference via the QVAC SDK, ensuring patient data remains on the device.

Ideal Use Cases

  • Research: Excellent for developers and researchers exploring medical language understanding and reasoning.
  • Application Development: Suitable for building prototypes and developer tools for health-related applications.
  • On-Device Inference: Optimized for privacy-sensitive medical information retrieval in edge environments.