Saxo/Linkbricks-Horizon-AI-Korean-Gemma-2-sft-dpo-27B

TEXT GENERATIONPricing:Input $2.6 / Cached $0.13 / Output $2.6Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kPublished:Aug 7, 2024License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Saxo/Linkbricks-Horizon-AI-Korean-Gemma-2-sft-dpo-27B is a Korean-specialized large language model developed by Yunsung Ji (Saxo) of Linkbricks Horizon-AI, based on the Gemma-2-27B-IT architecture. Fine-tuned with SFT and DPO methods, it leverages cross-lingual data (Korean, Chinese, English, Japanese) and logical datasets to enhance complex Korean logical reasoning and cross-lingual knowledge. This model excels in high-dimensional analysis of customer reviews, social media posts, coding, and robust content safety applications, including the detection of harmful content and sensitive information.

Loading preview...

Model Overview

Saxo/Linkbricks-Horizon-AI-Korean-Gemma-2-sft-dpo-27B is a specialized Korean large language model developed by Yunsung Ji (Saxo), a data scientist at Linkbricks Horizon-AI. It is built upon the Gemma-2-27B-IT base model and was fine-tuned using a two-stage SFT (Supervised Fine-Tuning) and DPO (Direct Preference Optimization) pipeline on NVIDIA H100 80GB GPUs.

Key Capabilities and Training

  • Multilingual and Logical Reasoning: The model was trained with cross-lingual datasets encompassing Korean, Chinese, English, and Japanese, alongside logical reasoning data. This approach enhances its ability to handle complex Korean logical problems and facilitates cross-lingual knowledge augmentation.
  • Content Safety and Analysis: A primary focus of this model is its strengthened capabilities in content safety. It is optimized for:
    • High-dimensional analysis of customer reviews and social media posts.
    • Coding tasks.
    • Detection of harmful content, including profanity, harassment, sexually explicit material, racist content, and other detrimental forms.
    • Enhanced detection of Personally Identifiable Information (PII) and sensitive information patterns.
    • Identification of high-risk requests that may require blocking or user warnings.
  • Tokenizer: It retains the original Gemma-2-27B-IT tokenizer without any vocabulary expansion.

Performance

The model achieved Rank-1 on the Open Ko LLM Leaderboard Season 2 (2024/11/01~2024/12/28) with an average score of 51.37. Notable benchmark scores include:

  • Ko-Winogrande: 68.27
  • Ko-GSM8k: 70.96
  • Ko-Harmlessness: 65.66

Technical Details

The fine-tuning process utilized Deepspeed Stage 3, rslora, and BAdam Layer Mode.