KETI-NLP/Qwen3.5-KETI-HAECHI-27B
Qwen3.5-KETI-HAECHI-27B is a 27 billion parameter multimodal model developed by KETI-NLP, derived from Qwen/Qwen3.5-27B, with a 32768 token context length. It is specifically optimized for Korean cultural-heritage understanding, Korean OCR, and enhanced tool calling for multi-step, stateful agent execution. The model significantly improves performance in these specialized areas while retaining general multimodal, language, and coding capabilities.
Loading preview...
Overview
Qwen3.5-KETI-HAECHI-27B is a 27 billion parameter multimodal model developed by KETI-NLP, building upon the Qwen/Qwen3.5-27B base. Its primary focus is to significantly enhance capabilities in Korean cultural-heritage understanding, Korean Optical Character Recognition (OCR), and robust tool calling for complex, multi-step agent execution. The model maintains the broad multimodal, language, and coding functionalities of its base.
Key Capabilities
- Korean Cultural Heritage & OCR: Excels at identifying official names of heritage objects, answering questions based on heritage images, and accurately reading Korean text from various sources including signs, scenes, and public documents. Qualitative comparisons show substantial improvements over the base model in recognizing specific heritage items and diverse Korean fonts and outdoor text.
- Tool Calling & Long-Horizon Task Execution: Designed to select and call tools effectively, carry information across multiple turns, track changing states, and work towards end-to-end goals over several steps. Benchmarks like Tau2 show an overall improvement in multi-step workflows, particularly in airline and retail domains.
- General Multimodal Understanding: Retains strong general multimodal interpretation, following both Korean and English instructions, and performing visual reasoning beyond its specialized domains. Quantitative benchmarks demonstrate significant gains in various general multimodal and language retention metrics, such as MMBench DEV EN v1.1 (+60.22 pp) and MMStar (+37.93 pp).
Good For
- Research and prototyping in Korean cultural-heritage image identification and visual question answering.
- Applications requiring robust Korean OCR for scene, sign, font, and public-document text.
- Developing multi-step, stateful tool-use workflows that require explicit monitoring and recovery mechanisms.
- Analysis of domain specialization and capability retention in multimodal models.
- General image understanding and text generation tasks, leveraging its strong base model capabilities.