SeongeonKim/gemma-2-9b-HangulFixer_v0.0

TEXT GENERATIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:16kPublished:Jan 24, 2025License:cc-by-nc-4.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

SeongeonKim/gemma-2-9b-HangulFixer_v0.0 is a 9 billion parameter causal language model developed by SeongeonKim, fine-tuned from unsloth/gemma-2-9b-bnb-4bit. This model is specifically designed and optimized for restoring obfuscated Korean hotel reviews to their original, clear, and natural form. It excels at text restoration tasks for Korean language inputs, leveraging a 16384 token context length.

Loading preview...

Model Overview

gemma-2-9b-HangulFixer is a 9 billion parameter text generation model developed by SeongeonKim, fine-tuned using Unsloth and Hugging Face's TRL library. It is based on the unsloth/gemma-2-9b-bnb-4bit model and is specifically engineered for the Korean language.

Key Capabilities

  • Korean Text Restoration: The primary function of this model is to de-obfuscate Korean hotel reviews that have been intentionally scrambled. This addresses a unique challenge where users obfuscate reviews to provide honest feedback without fear of deletion.
  • Specialized Training: It was fine-tuned on the SeongeonKim/ko-scrambled_v0.1 dataset, comprising 11,263 pairs of obfuscated and original Korean hotel reviews.
  • Efficiency: Training was accelerated using Unsloth, resulting in 2x faster completion compared to standard methods.

Use Cases and Limitations

This model is ideal for applications requiring the restoration of garbled Korean text, particularly in the context of user-generated content like reviews. It aims to improve communication between guests and accommodation providers by clarifying feedback.

Important Licensing Note: The model operates under the CC BY-NC 4.0 license, which permits non-commercial use only. Commercial applications require separate authorization, and proper attribution to the dataset source and license is mandatory for research use.