Markr-AI/Gukbap-Qwen2.5-7B
Markr-AI/Gukbap-Qwen2.5-7B is a 7.6 billion parameter Korean language model, fine-tuned from Qwen/Qwen2.5-7B-Instruct by HumanF-MarkrAI. This model is notable for being trained exclusively on a proprietary dataset generated using open-source models, avoiding data derived from private LLMs. It achieves a SOTA score of 8.39 on the LogicKor evaluation for Korean models under 7B parameters, demonstrating strong performance in reasoning, math, and writing tasks.
Loading preview...
Gukbap-Qwen2.5-7B: A Korean Language Model with Open-Source Data Integrity
Markr-AI/Gukbap-Qwen2.5-7B is a 7.6 billion parameter Korean language model, fine-tuned from Qwen/Qwen2.5-7B-Instruct. Developed by HumanF-MarkrAI, this model distinguishes itself by being trained entirely on a proprietary dataset generated exclusively through open-source models, thereby avoiding potential terms of service violations associated with using data from private LLMs like GPT-4.
Key Capabilities & Differentiators
- SOTA Korean Performance: Achieved an impressive score of 8.39 on the LogicKor evaluation, making it the state-of-the-art Korean language model under 7 billion parameters.
- Open-Source Data Integrity: Trained using datasets created solely with open-source LLMs (specifically
microsoft/WizardLM-2-8x22B), following methodologies from LIMA and WizardLM. - Strong Benchmarks: Demonstrates high scores in Korean reasoning (8.57), math (8.93), writing (9.50), and understanding (9.21) on the LogicKor benchmark, outperforming other 7B-class Korean models and even some larger models in specific categories.
Ideal Use Cases
- Korean Language Applications: Excellent for tasks requiring high proficiency in Korean, including complex reasoning, mathematical problem-solving, and creative writing.
- Ethical AI Development: Suitable for developers and organizations committed to building LLMs without relying on data generated by proprietary models, ensuring compliance with various terms of service.
- Research in Open-Source LLM Training: Provides a strong example of achieving top-tier performance using only open-source data generation and training methodologies.