ThakiCloud/Qwen3.8-27B-Satoori-KO
ThakiCloud/Qwen3.8-27B-Satoori-KO is a 27 billion parameter Qwen-based model developed by ThakiCloud, specifically designed for Korean regional dialect transformation. It covers five major Korean dialects (Gangwon, Gyeongsang, Jeolla, Chungcheong, and Jeju) and is trained directly on real dialect corpora. This model excels at single-turn tasks including dialect comprehension (transforming dialect to standard Korean), dialect identification, and dialect generation (transforming standard Korean to a specified dialect). It serves as a real-data reference for dialect transformation, achieving 80.1 chrF for comprehension and 79.4% accuracy for identification on KoDialectBench v0.
Loading preview...
Model Overview
ThakiCloud/Qwen3.8-27B-Satoori-KO is a 27 billion parameter model focused on Korean regional dialect transformation. It is trained directly on real dialect corpora for five regions: Gangwon, Gyeongsang, Jeolla, Chungcheong, and Jeju. This model acts as a real-data reference arm for dialect processing, designed to be compared with its synthetic-data counterpart, Qwen3.8-27B-Satoori-KO-Synth.
Key Capabilities
This model performs three single-turn tasks:
- Comprehension: Transforms a dialect sentence into its standard-Korean equivalent.
- Identification: Determines the regional origin of a given dialect sentence.
- Generation: Converts a standard Korean sentence into a specified regional dialect.
Performance Highlights
On KoDialectBench v0, this model demonstrates strong performance:
- Comprehension: Achieves 80.1 chrF and 24.7% exact match.
- Identification: Reaches 79.4% accuracy.
- Generation: Shows a dialectness score of 0.141 and 0.593 for region match.
Important Considerations
- Not a conversational model: It is trained exclusively on single-turn transformation and classification tasks, and therefore cannot hold natural conversations in dialect.
- Limitations: The dialectness metric is a lower bound, as it only counts marker lexicon words. The model processes transcribed text only, lacking audio-level pronunciation information. It was trained with a single seed, and its performance on other model scales is unverified.
Use Cases
This model is ideal for applications requiring precise, single-turn Korean dialect processing, such as:
- Translating dialect to standard Korean for better understanding.
- Identifying the regional origin of Korean text.
- Generating region-specific Korean dialect from standard input.