sionic-ai/PepperOCR-VL
PepperOCR-VL by Sionic AI is a 4.5 billion parameter vision-language model, fine-tuned from Qwen3.5-4B, designed to convert page images into structured Markdown. It excels at preserving text, headings, lists, tables (as HTML), and mathematical formulas (as LaTeX) from scanned documents in 17 languages. This model is optimized for document parsing, achieving state-of-the-art scores on MDPBench for Korean (92.6) and Thai (83.3).
Loading preview...
PepperOCR-VL: Multilingual Document Parsing to Markdown
PepperOCR-VL, developed by Sionic AI, is a 4.5 billion parameter vision-language model built upon Qwen3.5-4B. Its primary function is to transform diverse page images—including scanned documents, invoices, papers, and textbook pages—into clean, structured Markdown. The model accurately preserves layout elements such as headings, paragraphs, bullet/numbered lists, tables (rendered as HTML), and mathematical expressions (converted to LaTeX).
Key Capabilities
- End-to-End Document Parsing: Converts a single page image into a comprehensive Markdown output, maintaining reading order and structural integrity.
- Multilingual Support: Fine-tuned for 17 languages across Latin and non-Latin scripts, including German, English, Spanish, French, Arabic, Hindi, Japanese, Korean, Russian, Thai, and Chinese.
- Structured Output: Generates HTML
<table>blocks for tables and LaTeX for mathematical formulas ($x$for inline,$$x$$for display math). - High Performance: Achieves state-of-the-art scores on the MDPBench benchmark, with 92.6 for Korean and 83.3 for Thai as of September 2026.
Good For
- Agentic Document Parsing: Ideal for systems requiring structured extraction from various document types.
- Large-Scale Document Processing: Suitable for pipelines that need to convert vast quantities of scanned or photographed documents into machine-readable formats.
- Commercial Applications: Designed for integration into commercial products requiring robust OCR and document understanding capabilities.
- Research and Development: Provides a strong baseline for further fine-tuning or integration into multimodal research projects, particularly for complex document layouts and multilingual content.