Saxo/Linkbricks-Horizon-AI-Korean-Gemma-2-sft-dpo-27B
Saxo/Linkbricks-Horizon-AI-Korean-Gemma-2-sft-dpo-27B is a Korean-specialized large language model developed by Yunsung Ji (Saxo) of Linkbricks Horizon-AI, based on the Gemma-2-27B-IT architecture. Fine-tuned with SFT and DPO methods, it leverages cross-lingual data (Korean, Chinese, English, Japanese) and logical datasets to enhance complex Korean logical reasoning and cross-lingual knowledge. This model excels in high-dimensional analysis of customer reviews, social media posts, coding, and robust content safety applications, including the detection of harmful content and sensitive information.
Loading preview...
Model Overview
Saxo/Linkbricks-Horizon-AI-Korean-Gemma-2-sft-dpo-27B is a specialized Korean large language model developed by Yunsung Ji (Saxo), a data scientist at Linkbricks Horizon-AI. It is built upon the Gemma-2-27B-IT base model and was fine-tuned using a two-stage SFT (Supervised Fine-Tuning) and DPO (Direct Preference Optimization) pipeline on NVIDIA H100 80GB GPUs.
Key Capabilities and Training
- Multilingual and Logical Reasoning: The model was trained with cross-lingual datasets encompassing Korean, Chinese, English, and Japanese, alongside logical reasoning data. This approach enhances its ability to handle complex Korean logical problems and facilitates cross-lingual knowledge augmentation.
- Content Safety and Analysis: A primary focus of this model is its strengthened capabilities in content safety. It is optimized for:
- High-dimensional analysis of customer reviews and social media posts.
- Coding tasks.
- Detection of harmful content, including profanity, harassment, sexually explicit material, racist content, and other detrimental forms.
- Enhanced detection of Personally Identifiable Information (PII) and sensitive information patterns.
- Identification of high-risk requests that may require blocking or user warnings.
- Tokenizer: It retains the original Gemma-2-27B-IT tokenizer without any vocabulary expansion.
Performance
The model achieved Rank-1 on the Open Ko LLM Leaderboard Season 2 (2024/11/01~2024/12/28) with an average score of 51.37. Notable benchmark scores include:
- Ko-Winogrande: 68.27
- Ko-GSM8k: 70.96
- Ko-Harmlessness: 65.66
Technical Details
The fine-tuning process utilized Deepspeed Stage 3, rslora, and BAdam Layer Mode.