mkd-ai/Keural-Nova-v1.1
Keural Nova v1.1 by mkd-ai is a 35B-parameter Mixture-of-Experts (MoE) model, based on Qwen/Qwen3.6-35B-A3B, with approximately 3B active parameters per token. This instruction-tuned model is specifically optimized for high-quality Korean conversational AI and text comprehension, offering significant improvements in Korean language tasks. It supports a context length of up to 8,192 tokens and is intended for bilingual (Korean/English) assistant use and RAG-style grounded answering in Korean.
Loading preview...
Keural Nova v1.1: Korean-Focused Instruction-Tuned MoE Model
Keural Nova v1.1, developed by mkd-ai, is an instruction-tuned model built upon the Qwen/Qwen3.6-35B-A3B base architecture. This 35B-parameter Mixture-of-Experts (MoE) model, with ~3B active parameters per token, has been fine-tuned to excel in Korean conversational quality and text comprehension, while maintaining some English capabilities.
Key Capabilities & Performance
- Korean Language Proficiency: Substantially improves Korean conversational AI and text comprehension, as evidenced by a +4.41 score increase on the KoBEST benchmark compared to its base model.
- Bilingual Support: Designed for effective bilingual (Korean/English) assistant use.
- Context Window: Supports sequences up to 8,192 tokens, inheriting the base model's 262,144-token architectural context window (though long context behavior is untested).
- Robust Formatting: Demonstrates significant improvements in answer formatting, particularly for tasks like GSM8K.
Intended Use Cases
- Korean Conversational AI: Ideal for chatbots and virtual assistants requiring high-quality Korean dialogue.
- Korean Text Processing: Excellent for comprehension and generation of Korean text.
- RAG Systems: Suitable for RAG-style grounded answering in Korean.
Important Considerations & Limitations
- Not for Code Generation: This model is explicitly not suitable for code generation, showing a significant regression in HumanEval performance due to corrupted formatting in training data. Users needing code generation should use the base Qwen model.
- Knowledge-Heavy Tasks: Exhibits slight regressions (1-2 points) on general knowledge-recall benchmarks (KMMLU, HAE-RAE, MMLU) compared to the base model.
- Text-Only Fine-tune: While the base architecture is multimodal, this fine-tune focused solely on text; vision capabilities were neither trained nor evaluated.