ThakiCloud/Qwen3.8-27B-Satoori-KO-Synth

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 10, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

ThakiCloud/Qwen3.8-27B-Satoori-KO-Synth is a 27 billion parameter Qwen3.8 model developed by ThakiCloud, specifically designed for Korean regional dialect transformation. This model was trained exclusively on synthetic data, converting colloquial Korean utterances to various dialects using rule-based transformers, without direct exposure to source dialect text. It excels at transforming standard Korean into regional dialects and identifying dialect regions, offering a unique approach to dialect processing.

Loading preview...

Overview

ThakiCloud/Qwen3.8-27B-Satoori-KO-Synth is a 27 billion parameter Qwen3.8 model focused on Korean regional dialect transformation. Uniquely, it was trained entirely on synthetic data, where colloquial Korean sentences were converted into dialects using rule-based transformers, bypassing the use of actual dialect source text. This model serves as a comparative study to its real-data counterpart, exploring the extent of dialect processing achievable without restricted source data.

Key Capabilities

  • Korean Dialect Transformation: Converts standard Korean into various regional dialects (e.g., Gyeongsang, Jeolla, Jeju, Chungcheong, Gangwon).
  • Dialect Identification: Accurately identifies the region of a given dialect utterance.
  • Synthetic Data Training: Demonstrates a novel approach to training dialect models using only synthetically generated data.
  • Robustness to Formal Registers: Shows improved performance in maintaining regional differentiation for formal language registers compared to real-data models.

Performance Highlights

While not a conversational model, it achieves significant recovery rates compared to real-data models:

  • Identification: 91.2% recovery.
  • Generation (reference chrF): 72.4% recovery.
  • Axis-balanced mean: 75.8% recovery.

Limitations

  • Not a Conversational Model: Lacks conversational ability in dialect due to single-turn transformation training.
  • Regional Variance: Exhibits worse regional variance than the real-data model, particularly for Gangwon identification, and synthetic Chungcheong's marker ambiguity is higher.
  • Comprehension: Comprehension recovery remains around 52-54%.

Use Cases

  • Dialect Generation: Transforming standard Korean text into specific regional dialects.
  • Dialect Comprehension: Converting dialect text back into standard Korean.
  • Dialect Region Identification: Determining the geographical origin of a Korean dialect sentence.
  • Research: Ideal for researchers studying synthetic data generation, rule-based transformation, and the efficacy of training LLMs without direct access to sensitive or restricted datasets.