GL3MON/qwen3-vl-2b-medical-pmcvqa-sft-merged

VISIONPricing:Input $0.32 / Cached $0.016 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 18, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The GL3MON/qwen3-vl-2b-medical-pmcvqa-sft-merged is a 2 billion parameter Qwen3-VL model developed by GL3MON, specifically fine-tuned for medical visual question answering (VQA) tasks. This model leverages the Qwen3-VL architecture and was efficiently trained using Unsloth and Huggingface's TRL library. It is designed to process and understand medical images in conjunction with textual queries, making it suitable for specialized medical AI applications. The model has a context length of 32768 tokens.

Loading preview...

GL3MON/qwen3-vl-2b-medical-pmcvqa-sft-merged: Medical VQA Model

This model, developed by GL3MON, is a specialized 2 billion parameter Qwen3-VL variant fine-tuned for medical visual question answering (VQA). It builds upon the base GL3MON/qwen3-vl-2b-medical-sft model, enhancing its capabilities for medical domain-specific queries.

Key Capabilities

  • Medical Visual Question Answering (VQA): Designed to interpret medical images and answer questions related to their content, making it suitable for diagnostic support or educational tools.
  • Efficient Training: The model was fine-tuned using Unsloth and Huggingface's TRL library, enabling faster training times.
  • Qwen3-VL Architecture: Leverages the robust Qwen3-VL architecture, indicating strong multimodal understanding capabilities.
  • Context Length: Supports a substantial context length of 32768 tokens, allowing for processing of detailed visual and textual inputs.

Good For

  • Medical AI Applications: Ideal for use cases requiring the analysis of medical imagery combined with natural language questions.
  • Research and Development: Provides a specialized foundation for further research in medical multimodal AI.
  • Educational Tools: Can be integrated into systems that help medical students or professionals understand complex medical visuals.