zenlm/zen-vl-4b-agent

VISIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Nov 4, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

zenlm/zen-vl-4b-agent is a 4 billion parameter vision-language agent model, repackaged from Alibaba Qwen's Qwen3-VL-4B-Instruct. This model is designed for image understanding, OCR, and visual reasoning tasks. It supports multimodal inputs including text, image, and video, making it suitable for applications requiring comprehensive visual and textual analysis.

Loading preview...

zenlm/zen-vl-4b-agent Overview

This model, zenlm/zen-vl-4b-agent, is a 4 billion parameter vision-language agent. It is a repackaged version of Alibaba Qwen's Qwen/Qwen3-VL-4B-Instruct, distributed under an Apache-2.0 license for the OSS-clean Zen model line. It is important to note that this model was not trained from scratch by zenlm but rather provides a permissively-licensed redistribution of the original Qwen model.

Key Capabilities

  • Multimodal Understanding: Processes and integrates information from text, images, and video inputs.
  • Image Understanding: Capable of interpreting visual content within images.
  • Optical Character Recognition (OCR): Designed to extract text from images.
  • Visual Reasoning: Performs reasoning tasks based on visual information.

Good For

  • Applications requiring a vision-language model with a 4 billion parameter count.
  • Tasks involving comprehensive analysis of visual and textual data.
  • Use cases in image understanding, OCR, and visual reasoning where a Qwen3-VL architecture is suitable.

Note: This model has been superseded by zenlm/zen3-vl, which is the canonical name for the updated version. The weights for zen-vl-4b-agent remain available for reproducibility purposes.