nectec/Pathumma-llm-vision-3.0.0-preview

VISIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2.3BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 25, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Pathumma-llm-vision-3.0.0-preview is a 2.3 billion parameter vision-language model developed by NECTEC, based on Qwen3.5-2B. It is specifically optimized for Thai OCR and multilingual image-text understanding, trained on 377K OCR samples. This model excels at document parsing, text recognition, and cognitive visual question answering, making it suitable for efficient OCR and document understanding deployments.

Loading preview...

Model Overview

Pathumma-llm-vision-3.0.0-preview is a 2.3 billion parameter vision-language model developed by NECTEC. Built upon the Qwen3.5-2B architecture, this model has been extensively trained on 377K OCR samples to enhance its performance in Thai and multilingual OCR, as well as general image-text understanding.

Key Capabilities

  • Thai OCR Optimization: Specifically designed and trained to improve optical character recognition for Thai language content.
  • Multilingual Image-Text Understanding: Capable of processing and understanding text within images across multiple languages.
  • Document and Scene Text Recognition: Excels at extracting text from various real-world documents and scene images.
  • Efficient Deployment: Engineered for compact and efficient deployment in OCR and document understanding applications.

Performance Highlights

Evaluated on ThaiOCRBench, Pathumma-llm-vision-3.0.0-preview demonstrates strong performance, achieving an overall score of 0.5050. Notable strengths include:

  • Document Parsing: Achieved 0.4872.
  • Full-page OCR: Scored 0.7831.
  • Diagram VQA: Reached 0.5784.
  • Cognition VQA: Attained 0.6763.
  • Infographics: Scored 0.6451.

Intended Use Cases

This model is particularly well-suited for:

  • Thai OCR tasks.
  • Scene text recognition.
  • Document text extraction.
  • Thai document understanding applications.
  • Deployments requiring efficient OCR capabilities.