JEJUMA/JEJUMA-002

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 4, 2024License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

JEJUMA/JEJUMA-002 is an 8 billion parameter language model developed by JEJUMA, fine-tuned from Llama3.1 to specialize in Korean regional dialects. It excels at tasks such as converting dialects to standard Korean, standard Korean to dialects, and detecting dialect types. Trained on approximately 5 million dialect-to-standard Korean pair data, it aims to preserve and facilitate understanding of endangered regional dialects, particularly Jeju dialect.

Loading preview...

JEJUMA-002: Preserving Korean Regional Dialects

JEJUMA-002 is an 8 billion parameter language model, fine-tuned from Llama3.1, specifically designed to address the rapid disappearance of Korean regional dialects. Many dialects are becoming less common, even among younger generations, and existing language models struggle with them due to a lack of online data, especially for distinct dialects like Jeju-eo.

Key Capabilities

  • Dialect-to-Standard Conversion: Translates regional dialects into standard Korean.
  • Standard-to-Dialect Generation: Creates dialect versions from standard Korean.
  • Dialect Detection: Identifies the region of a given dialect.
  • Automated Detection and Conversion: Automatically detects a dialect and converts it to standard Korean.

Training and Performance

JEJUMA-002 was trained using approximately 5 million regional dialect-to-standard Korean pair datasets, with 750,000 carefully selected for strong dialect features. This data was expanded to 2 million entries across four tasks. The model was trained using LlamaFactory with LoRA for one epoch. It demonstrates high translation accuracy for difficult dialects, comparable to or exceeding models like GPT-4o, Upstage Solar, and Naver HCX, particularly for Jeju-eo.

Use Cases

This model is ideal for applications focused on:

  • Cultural Preservation: Helping to document and maintain endangered Korean dialects.
  • Language Education: Assisting learners in understanding and generating regional speech patterns.
  • Content Localization: Adapting standard Korean content into various regional dialects.