wwwcomcom/ko-morph-qwen2.5-0.5b-instruct

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The wwwcomcom/ko-morph-qwen2.5-0.5b-instruct is a 0.5 billion parameter instruction-tuned model based on the Qwen2.5 architecture, specifically fine-tuned for Korean language tasks. It utilizes a morphological tokenizer and is optimized for Korean grammar and general knowledge, demonstrating improved performance over its base model in these areas. This model is particularly suited for Korean natural language understanding and generation tasks where a compact model size is beneficial.

Loading preview...

Overview

This model, ko-morph-qwen2.5-0.5b-instruct, is a 0.5 billion parameter instruction-tuned variant of the Qwen2.5 architecture, developed by wwwcomcom. It was fine-tuned using the beomi/KoAlpaca-v1.1a dataset, which consists of Korean knowledge-based questions and GPT-generated answers.

Key Characteristics

  • Morphological Tokenizer: Employs a morphological tokenizer, which influences its performance, particularly in Korean grammar tasks.
  • Instruction-Tuned: Specifically trained with a ### 질문: {instruction} ### 답변: {output} prompt format, without newlines, due to tokenizer limitations.
  • Korean Language Optimization: Shows significant improvements in Korean grammar and general knowledge compared to the original Qwen2.5-0.5B base model under identical SFT conditions.

Performance Highlights

Evaluations against a Qwen2.5-0.5B base model, under the same SFT conditions, reveal:

  • Grammar: Achieved 97.9% (327/334) on grammar minimal pairs, significantly outperforming the Qwen original's 87.7% (293/334).
  • Korean Knowledge: Scored 3 accurate and 3 partial answers out of 8 for Korean knowledge/common sense, whereas the Qwen original had 8 incorrect answers.
  • Limitations: Arithmetic and formal logic reasoning are noted to be better in the original Qwen model, suggesting potential catastrophic forgetting due to the Korean web text CPT on top of Qwen's math/code pre-training.

Usage Considerations

  • Prompt Format: Strictly adhere to the ### 질문: {instruction} ### 답변: format without newlines, as the model was trained this way.
  • No ChatML: Avoid ChatML or prompts with newlines, as they can degrade performance.

Licensing

The model follows the Apache 2.0 license, inherited from its Qwen2.5-0.5B base.