leoyinn/qwen3vl-flare25

VISIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Nov 13, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

leoyinn/qwen3vl-flare25 is a 4 billion parameter vision-language model, fully fine-tuned from Qwen/Qwen3-VL-4B-Instruct by leoyinn. It specializes in medical imaging tasks, trained on the FLARE 2025 dataset across 8 modalities and 19 datasets. This model excels at classification, detection, regression, counting, and report generation for medical images, offering significant performance improvements over its baseline.

Loading preview...

Model Overview

This model, leoyinn/qwen3vl-flare25, is a specialized 4 billion parameter vision-language model derived from the Qwen3-VL-4B-Instruct architecture. It has been extensively fine-tuned on the FLARE 2025 medical imaging dataset, designed to address a wide array of medical vision-language tasks.

Key Capabilities

  • Multi-modal Medical Imaging: Supports 8 distinct medical imaging modalities, including Ultrasound, X-ray, Retinography, Microscopy, Clinical Photography, Dermatology, Endoscopy, and Mammography.
  • Diverse Task Performance: Capable of performing various tasks such as Classification, Multi-label Classification, Detection, Instance Detection, Regression, Counting, and Report Generation.
  • Significant Performance Gains: Demonstrates substantial improvements over the baseline model, with classification balanced accuracy increasing by over 2,300% and detection becoming a new capability with 80.3% [email protected].
  • Extensive Training Data: Trained on 12,232 samples across 26 datasets, ensuring broad applicability within medical imaging.

Good For

  • Medical Image Analysis: Ideal for applications requiring automated analysis of medical images across various modalities.
  • Clinical Decision Support: Can be integrated into systems for tasks like disease classification, anomaly detection, and automated report generation.
  • Research and Development: Provides a strong foundation for further research in medical AI, particularly in multi-modal learning and specialized medical vision-language models.