Markr-AI/Gukbap-Qwen2.5-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 25, 2024Architecture:Transformer0.0K Featherless Exclusive Cold

Markr-AI/Gukbap-Qwen2.5-7B is a 7.6 billion parameter Korean language model, fine-tuned from Qwen/Qwen2.5-7B-Instruct by HumanF-MarkrAI. This model is notable for being trained exclusively on a proprietary dataset generated using open-source models, avoiding data derived from private LLMs. It achieves a SOTA score of 8.39 on the LogicKor evaluation for Korean models under 7B parameters, demonstrating strong performance in reasoning, math, and writing tasks.

Loading preview...

Gukbap-Qwen2.5-7B: A Korean Language Model with Open-Source Data Integrity

Markr-AI/Gukbap-Qwen2.5-7B is a 7.6 billion parameter Korean language model, fine-tuned from Qwen/Qwen2.5-7B-Instruct. Developed by HumanF-MarkrAI, this model distinguishes itself by being trained entirely on a proprietary dataset generated exclusively through open-source models, thereby avoiding potential terms of service violations associated with using data from private LLMs like GPT-4.

Key Capabilities & Differentiators

  • SOTA Korean Performance: Achieved an impressive score of 8.39 on the LogicKor evaluation, making it the state-of-the-art Korean language model under 7 billion parameters.
  • Open-Source Data Integrity: Trained using datasets created solely with open-source LLMs (specifically microsoft/WizardLM-2-8x22B), following methodologies from LIMA and WizardLM.
  • Strong Benchmarks: Demonstrates high scores in Korean reasoning (8.57), math (8.93), writing (9.50), and understanding (9.21) on the LogicKor benchmark, outperforming other 7B-class Korean models and even some larger models in specific categories.

Ideal Use Cases

  • Korean Language Applications: Excellent for tasks requiring high proficiency in Korean, including complex reasoning, mathematical problem-solving, and creative writing.
  • Ethical AI Development: Suitable for developers and organizations committed to building LLMs without relying on data generated by proprietary models, ensuring compliance with various terms of service.
  • Research in Open-Source LLM Training: Provides a strong example of achieving top-tier performance using only open-source data generation and training methodologies.