datalab-to/surya-ocr-2
Surya OCR 2 by Datalab is a 0.65 billion parameter vision-language model (VLM) designed for document intelligence, offering comprehensive OCR, layout analysis, and table recognition. It achieves 83.3% on the olmOCR-bench benchmark, making it a top performer under 3 billion parameters, and boasts multilingual support across 91 languages with an 87.2% overall pass rate. This model is optimized for high-speed document processing, capable of 5 pages/second throughput on an RTX 5090, making it suitable for efficient and accurate extraction of structured data from diverse document types.
Loading preview...
Surya OCR 2: Document Intelligence VLM
Surya OCR 2, developed by Datalab, is a 0.65 billion parameter vision-language model (VLM) specifically engineered for advanced document intelligence tasks. It integrates OCR, layout analysis, and table recognition into a single model, providing a comprehensive solution for extracting information from various document types.
Key Capabilities
- High Accuracy: Achieves 83.3% on the olmOCR-bench benchmark, positioning it as a leading model under 3 billion parameters for document parsing.
- Multilingual Support: Scores 87.2% across an internal benchmark of 91 languages, with strong performance in widely spoken languages like English (92.3%), Spanish (90.7%), and German (89.7%).
- Speed: Delivers a throughput of 5 pages per second on an RTX 5090, making it highly efficient for large-scale document processing.
- Layout Analysis: Provides detailed layout analysis, identifying elements such as tables, images, headers, and text blocks, along with reading order.
- Table Recognition: Accurately recognizes table structures, including rows and columns, and can output full HTML tables.
- Math/Equations: Handles inline math and equations, outputting them in KaTeX-compatible LaTeX within HTML results.
Good For
- Automated Document Processing: Ideal for applications requiring fast and accurate extraction of text, layout, and tabular data from scanned documents, PDFs, and images.
- Multilingual Data Extraction: Suitable for global applications needing to process documents in a wide array of languages.
- Structured Data Capture: Excels in use cases where precise identification of document structure, including complex tables and reading order, is critical.