ut-amrl/foresight-qwen3vl-2b-sft
The ut-amrl/foresight-qwen3vl-2b-sft is a 2 billion parameter supervised fine-tuned vision-language model based on the Qwen3-VL architecture, developed by UT-AMRL. It is specifically designed for robotic navigation, enabling a policy to iteratively identify instruction-relevant visual cues and refine motion plans. This model proposes image-space trajectories and self-critiques proposals, serving both motion planning and critique roles within a single vLLM engine.
Loading preview...
Model Overview
The ut-amrl/foresight-qwen3vl-2b-sft is a 2 billion parameter supervised fine-tuned (SFT) vision-language model (VLM) developed by UT-AMRL. It is built upon the Qwen/Qwen3-VL-2B-Instruct base model and is a core component of the Foresight navigation policy. This model is engineered to process natural-language goals and RGB observations to propose and critique image-space trajectories for open-world navigation.
Key Capabilities
- Vision-Language Integration: Combines visual input with natural language instructions to generate navigation plans.
- Trajectory Proposal: Given a goal and visual history, it proposes an image-space trajectory.
- Self-Critique: The model is co-trained to critique its own proposed motion plans, allowing for iterative refinement.
- Efficient Deployment: Both planning and critique functionalities are integrated into a single checkpoint, suitable for deployment via vLLM.
- Fine-tuned Performance: Fine-tuned from
Qwen/Qwen3-VL-2B-Instructusing LoRA (rank 64, alpha 64) across language and vision towers, with merged adapters for inference without PEFT dependency.
Good For
- Robotic Navigation: Ideal for research and development in autonomous navigation systems that require iterative reasoning and visual clue identification.
- Embodied AI: Suitable for applications where an agent needs to understand and act upon visual and linguistic instructions in dynamic environments.
- VLM-based Control: Provides a foundation for developing control policies that leverage advanced vision-language understanding.