zr-wang/AgenticOCR-4B
VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 4, 2026Architecture:Transformer0.0K Featherless Exclusive Cold
zr-wang/AgenticOCR-4B is a 4 billion parameter Qwen3-VL-4B-Instruct model, developed by zr-wang, specifically trained for the AgenticOCR project. This model is designed for query-driven, on-demand visual evidence extraction from documents, supporting Retrieval Augmented Generation (RAG) workflows. It integrates visual processing capabilities with a 32768 token context length to enhance document understanding and information retrieval.
Loading preview...
AgenticOCR-4B Overview
AgenticOCR-4B is a specialized 4 billion parameter vision-language model, built upon the Qwen3-VL-4B-Instruct architecture. Developed by zr-wang, its primary purpose is to facilitate the AgenticOCR project, focusing on advanced document processing.
Key Capabilities
- Query-Driven Visual Evidence Extraction: The model excels at identifying and extracting specific visual information from documents based on user queries.
- On-Demand Information Retrieval: It supports dynamic retrieval of visual evidence, crucial for enhancing Retrieval Augmented Generation (RAG) systems.
- Document RAG Enhancement: By providing precise visual context, AgenticOCR-4B improves the accuracy and relevance of information retrieved for RAG applications.
- Integrated Visual Processing: As a vision-language model, it processes both textual and visual inputs, enabling a deeper understanding of document content.
- Extended Context Length: Features a substantial context length of 32768 tokens, allowing for comprehensive analysis of longer documents.
Good For
- Document Analysis: Ideal for tasks requiring detailed understanding and extraction of information from complex documents.
- RAG Systems: Particularly suited for developers building RAG applications that benefit from visual evidence.
- Information Extraction: Useful for scenarios where specific data points, including those embedded in images or layouts, need to be precisely located and extracted based on queries.