xCloudinfo/Gemma-4-E4B-xReadable-zhTW-OCR-v1

VISIONConcurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 25, 2026License:gemmaArchitecture:Transformer Featherless Exclusive Cold

xCloudinfo/Gemma-4-E4B-xReadable-zhTW-OCR-v1 is a 7.9 billion parameter Gemma 4-E4B-it based model developed by xCloudinfo, fine-tuned for Traditional Chinese (Taiwan) Optical Character Recognition (OCR). It excels at recognizing mixed Chinese and English content in documents, invoices, and notes, supporting on-premise deployment. This model offers improved character error rates over previous versions and external benchmarks for printed text, and can be deployed using both safetensors and GGUF formats.

Loading preview...

Overview

xCloudinfo/Gemma-4-E4B-xReadable-zhTW-OCR-v1 is a 7.9 billion parameter model based on Google's gemma-4-E4B-it architecture, fine-tuned by xCloudinfo for Traditional Chinese (Taiwan) OCR. It specializes in recognizing text from Taiwanese documents, including official papers, invoices, and notes, supporting both Chinese and English mixed content. A key differentiator is its ability to be deployed on-premise, ensuring data privacy.

Key Capabilities

  • Superior OCR Performance: Achieves significantly lower Character Error Rates (CER) compared to its Llama-based predecessor and external benchmarks like nvidia/nemotron-ocr-v2 for unseen fonts (Songti, STHeiti) and reserved font sets.
  • Multimodal Support: Built on Gemma 4 Vision, it inherits native vision capabilities. The visual encoder is frozen during fine-tuning, with only the language side trained via LoRA.
  • Flexible Deployment: Available in both bf16 safetensors for high-precision server inference (transformers/vLLM) and GGUF format for lightweight, on-premise deployment with llama.cpp, Ollama, or LM Studio, including vision support via mmproj-f16.
  • Traditional Chinese Normalization: All outputs are consistently normalized to Traditional Chinese.

Use Cases

  • Taiwanese Document Digitization: Ideal for digitizing printed or scanned documents in Traditional Chinese.
  • Invoice and Note Recognition: Suitable for pre-processing and recognizing text from invoices and handwritten notes.
  • On-premise OCR: Designed for scenarios requiring local deployment and data privacy.

Limitations

  • Handwriting Recognition: Still experimental, with CERs ranging from 40% for clear text to 64-88% for messy or faint handwriting, often limited by image quality rather than model capability. Manual review is recommended for critical applications.
  • Complex Layouts: For tables or complex page layouts, pre-processing like line or cell segmentation is advised.
  • Non-definitive Output: Outputs are for recognition reference and should not be the sole basis for critical decisions (e.g., legal, medical, financial).