ut-amrl/foresight-qwen3vl-2b-sft

VISIONPricing:Input $0.32 / Cached $0.016 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 11, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The ut-amrl/foresight-qwen3vl-2b-sft is a 2 billion parameter supervised fine-tuned vision-language model based on the Qwen3-VL architecture, developed by UT-AMRL. It is specifically designed for robotic navigation, enabling a policy to iteratively identify instruction-relevant visual cues and refine motion plans. This model proposes image-space trajectories and self-critiques proposals, serving both motion planning and critique roles within a single vLLM engine.

Loading preview...

Model Overview

The ut-amrl/foresight-qwen3vl-2b-sft is a 2 billion parameter supervised fine-tuned (SFT) vision-language model (VLM) developed by UT-AMRL. It is built upon the Qwen/Qwen3-VL-2B-Instruct base model and is a core component of the Foresight navigation policy. This model is engineered to process natural-language goals and RGB observations to propose and critique image-space trajectories for open-world navigation.

Key Capabilities

  • Vision-Language Integration: Combines visual input with natural language instructions to generate navigation plans.
  • Trajectory Proposal: Given a goal and visual history, it proposes an image-space trajectory.
  • Self-Critique: The model is co-trained to critique its own proposed motion plans, allowing for iterative refinement.
  • Efficient Deployment: Both planning and critique functionalities are integrated into a single checkpoint, suitable for deployment via vLLM.
  • Fine-tuned Performance: Fine-tuned from Qwen/Qwen3-VL-2B-Instruct using LoRA (rank 64, alpha 64) across language and vision towers, with merged adapters for inference without PEFT dependency.

Good For

  • Robotic Navigation: Ideal for research and development in autonomous navigation systems that require iterative reasoning and visual clue identification.
  • Embodied AI: Suitable for applications where an agent needs to understand and act upon visual and linguistic instructions in dynamic environments.
  • VLM-based Control: Provides a foundation for developing control policies that leverage advanced vision-language understanding.