MirilAI/Miril-DroneVLM-2B-2

VISIONConcurrent Unit Cost:1Model Size:5.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 24, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

MirilAI/Miril-DroneVLM-2B-2 is a 5.1 billion parameter open-weight aerial vision-language model developed by Miril.ai. It is designed to enable drones to reason about their environment by processing overhead images and English questions, returning machine-readable JSON responses for tasks like scene captioning, factual answering, and precise object pointing. This model excels at providing structured outputs for civilian drone applications such as delivery, infrastructure inspection, and disaster recovery, making it a perception component for easier system supervision and integration.

Loading preview...

Miril-DroneVLM-2B-2: Aerial Vision-Language Model

Miril-DroneVLM-2B-2 is a 5.1 billion parameter open-weight aerial vision-language model from Miril.ai, designed to enable drones to interpret and respond to their environment. It processes overhead images and natural English questions, returning one of four machine-readable JSON responses: a whole-scene caption, a factual answer, a landing/delivery location, or a point on a visible target. This unique approach allows downstream software to deterministically dispatch actions based on the model's typed output, rather than parsing free-form text.

Key Capabilities

  • Structured JSON Output: Provides caption, answer, location, or pointing JSON responses based on user intent, simplifying integration with automated systems.
  • Aerial Image Understanding: Specialized for drone-view imagery, supporting applications like infrastructure inspection, agriculture, and disaster recovery.
  • Spatial Reasoning: Can identify and provide coordinates for specific objects or suggest broad directional cues for landing/delivery.
  • Edge-Oriented Design: A 2B-class model suitable for testing and integration into edge-based drone systems.
  • Experimental Spoken Input: Inherits Gemma's audio path, allowing for experimental zero-shot spoken question inference, though fine-tuned primarily on typed questions.

Good for

  • Automated Drone Operations: Ideal for prototyping and research in civilian drone applications requiring structured environmental understanding.
  • Situational Awareness: Enhancing first-response and disaster-recovery efforts by quickly summarizing aerial imagery.
  • Inspection Triage: Streamlining human review processes for infrastructure, agriculture, and construction by localizing points of interest.
  • Physical AI Experiments: Exploring typed-output interfaces for physical AI systems where precise, machine-readable responses are critical.