ThakiCloud/Qwen3.8-27B-Human-KO-Safety
ThakiCloud/Qwen3.8-27B-Human-KO-Safety is a 27 billion parameter Qwen3.8-27B adaptation by ThakiCloud, fine-tuned for Korean conversation with a focus on safety and bias mitigation. This model utilizes preference learning (DPO) to achieve a 92.8% withheld response rate on ambiguous social bias questions, significantly reducing stereotypical answers. It maintains strong performance across benchmarks like HumanEval (96.0%) and MMLU English (93.0%), while improving Korean human-likeness and KMMLU scores. The model is designed to provide safer, less biased responses in Korean contexts by abstaining from ambiguous questions.
Loading preview...
Overview
ThakiCloud/Qwen3.8-27B-Human-KO-Safety is a specialized 27 billion parameter model derived from Qwen/Qwen3.8-27B, adapted for Korean conversational safety. It integrates style alignment (Human-KO) with preference learning (DPO) to enhance its ability to withhold answers on ambiguous questions where social bias might be present, while still providing direct answers to well-defined queries.
Key Capabilities and Features
- Bias Mitigation: Achieves a 92.8% withheld response rate on ambiguous questions in the KoBBQ benchmark, significantly reducing stereotypical answers to 6.7% (compared to 28.0% for EXAONE reference).
- Robust Performance: Maintains high performance on core benchmarks, including 96.0% on HumanEval, 93.0% on MMLU English, and 96.0% on GPQA diamond.
- Korean Language Enhancement: Shows a notable +4.2pp improvement on KMMLU and an increased Korean human-likeness win rate.
- Defect Fix: Trained with
enable_thinking=Falseto prevent empty responses in thinking mode, a known defect in earlier DPO training prompts. - Context Length: Supports a substantial context length of 32768 tokens.
Good For
- Applications requiring high safety and bias awareness in Korean language interactions.
- Use cases where the model needs to responsibly abstain from answering ambiguous or potentially biased questions.
- Developers seeking a robust Korean-adapted LLM that balances performance with ethical response generation. Note that while it reduces stereotype answers, the conditional bias score among answered items is not less biased.