apple/LensVLM-9B

VISIONPricing:Input $0.1078 / Cached $0.0862 / Output $0.28Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026License:apple-amlrArchitecture:Transformer0.2K Featherless Exclusive Cold

LensVLM-9B is a 9 billion parameter Vision Language Model (VLM) developed by Apple that processes compressed visual representations of text. It uniquely scans compressed images of text and selectively expands only relevant pages to their uncompressed form using learned tools. This model is designed for efficient document understanding by focusing computational resources on pertinent sections.

Loading preview...

LensVLM-9B: Selective Context Expansion for Visual Text

LensVLM-9B, developed by Apple, is a 9 billion parameter Vision Language Model (VLM) designed for efficient processing of text within documents. Its core innovation lies in its ability to handle compressed visual representations of text, selectively expanding only the relevant pages to their uncompressed form through learned tools. This approach allows the model to focus computational effort on critical sections of a document, potentially improving efficiency and reducing processing overhead.

Key Capabilities

  • Selective Context Expansion: Processes compressed visual text and intelligently expands only pertinent pages.
  • Efficient Document Understanding: Optimizes resource allocation by focusing on relevant document sections.
  • Vision Language Integration: Combines visual processing with language understanding for document analysis.

Use Cases

  • Document Analysis: Ideal for tasks requiring understanding of large documents where only specific sections are relevant.
  • Information Extraction: Can be applied to extract information from visually presented text, especially in compressed formats.
  • Research and Development: Provides a foundation for further research into efficient VLM architectures and selective processing techniques.

For more technical details, the associated research paper, "LensVLM: Selective Context Expansion for Compressed Visual Representation of Text," is available on arXiv, and the model's code can be found on GitHub.