maravilla1/Qwen2.5-VL-3B-Instruct-fork

VISIONPricing:Input $0.32 / Cached $0.016 / Output $1.6Concurrent Unit Cost:1Model Size:3BQuant:BF16Context Size:32kPublished:Aug 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

maravilla1/Qwen2.5-VL-3B-Instruct-fork, named Mira-Sight, is a 3 billion parameter vision-language instruction-tuned model based on the Qwen2.5 architecture. Developed by Mira AI, this model is designed for multimodal tasks, specifically processing both image and text inputs to generate text outputs. Its primary differentiator is its vision capabilities, enabling it to understand and respond to queries involving visual information. This model is suitable for applications requiring image-to-text understanding and generation.

Loading preview...

Mira-Sight: A Vision-Language Instruction Model

Mira-Sight (maravilla1/Qwen2.5-VL-3B-Instruct-fork) is a 3 billion parameter instruction-tuned model developed by Mira AI, building upon the Qwen2.5 architecture. This model is specifically designed for multimodal applications, integrating both visual and textual understanding to produce relevant text outputs.

Key Capabilities

  • Vision-Language Integration: Processes both image and text inputs simultaneously.
  • Image-to-Text Generation: Capable of generating descriptive or analytical text based on provided images.
  • Instruction Following: Fine-tuned to follow instructions for various multimodal tasks.

Good For

  • Visual Question Answering: Answering questions about the content of images.
  • Image Captioning: Generating descriptive captions for images.
  • Multimodal Chatbots: Developing conversational agents that can interact with users using both text and visual information.