zr-wang/AgenticOCR-8B
VISIONPricing:Input $0.727 / Output $5.405Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 4, 2026Architecture:Transformer0.0K Featherless Exclusive Cold
AgenticOCR-8B is an 8 billion parameter Qwen3-VL-8B-Instruct model developed by zr-wang, fine-tuned for query-driven, on-demand visual evidence extraction in document RAG applications. With a context length of 32768 tokens, this model specializes in processing visual documents to retrieve specific information based on user queries. It is designed to enhance document understanding and retrieval by integrating visual and language capabilities.
Loading preview...
AgenticOCR-8B Overview
AgenticOCR-8B is an 8 billion parameter model based on the Qwen3-VL-8B-Instruct architecture, developed by zr-wang. It is specifically trained for the AgenticOCR project, focusing on advanced document understanding and information extraction.
Key Capabilities
- Query-Driven Visual Evidence Extraction: The model excels at identifying and extracting specific visual evidence from documents based on natural language queries.
- On-Demand Information Retrieval: Designed for dynamic retrieval of information, making it suitable for interactive document analysis.
- Document RAG Enhancement: Integrates visual processing into Retrieval Augmented Generation (RAG) workflows, improving the accuracy and relevance of retrieved information from documents.
- Visual Language Understanding: Leverages the Qwen3-VL foundation to process both textual and visual elements within documents.
Good For
- Applications requiring precise information extraction from scanned documents or images.
- Building intelligent document processing systems that can answer questions based on visual content.
- Enhancing RAG pipelines where visual context is crucial for accurate retrieval.
- Research and development in multimodal AI for document analysis.