zenlm/zen-vl-8b-instruct

VISIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Nov 4, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The zenlm/zen-vl-8b-instruct is an 8 billion parameter vision-language model, repackaged from Alibaba Qwen's Qwen3-VL-8B-Instruct. This model is designed for image understanding, optical character recognition (OCR), and visual reasoning tasks. It supports text, image, and video modalities, making it suitable for multimodal applications requiring robust visual comprehension.

Loading preview...

zen-vl-8b-instruct Overview

zen-vl-8b-instruct is an 8 billion parameter vision-language model, based on the Qwen3-VL architecture. It is a permissively-licensed redistribution of Alibaba Qwen's Qwen3-VL-8B-Instruct, intended for the OSS-clean Zen model line. This model is not trained from scratch but provides a readily available, open-source option for multimodal AI development.

Key Capabilities

  • Multimodal Understanding: Processes and interprets information from text, images, and video inputs.
  • Image Understanding: Excels at comprehending visual content within images.
  • Optical Character Recognition (OCR): Capable of extracting text from images.
  • Visual Reasoning: Performs tasks that require logical inference based on visual data.

Good For

  • Applications requiring image analysis and visual question answering.
  • Developing systems that need to extract text from diverse visual sources.
  • Use cases involving multimodal data processing where both visual and textual context are crucial.

Note: This model has been superseded by zenlm/zen3-vl, which is the canonical name for the latest version. The weights for zen-vl-8b-instruct remain available for reproducibility purposes.