tayaee/Qwen2.5-1.5B-Korean-DPO-smoke
The tayaee/Qwen2.5-1.5B-Korean-DPO-smoke is a 1.54 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen Team. This model features a transformer architecture with a 32,768-token context length, offering enhanced capabilities in coding, mathematics, and long-text generation. It is specifically fine-tuned for instruction following, structured data understanding, and multilingual support across 29 languages, including Korean.
Loading preview...
Qwen2.5-1.5B-Korean-DPO-smoke Overview
This model is an instruction-tuned variant of the Qwen2.5 series, a family of large language models developed by the Qwen Team. With 1.54 billion parameters and a 32,768-token context length, it builds upon the Qwen2 architecture with significant improvements.
Key Capabilities
- Enhanced Knowledge & Reasoning: Demonstrates improved capabilities in coding and mathematics, leveraging specialized expert models.
- Instruction Following: Shows significant advancements in adhering to instructions and generating diverse outputs.
- Long-Context Generation: Capable of generating long texts (over 8K tokens) and understanding structured data like tables.
- Structured Output: Excels at generating structured outputs, particularly JSON, and is resilient to varied system prompts for role-play.
- Multilingual Support: Supports over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, and Korean.
When to Use This Model
This model is particularly well-suited for applications requiring robust instruction following, generation of structured data, and handling long-context conversations. Its multilingual capabilities make it valuable for Korean-specific tasks, while its improved coding and mathematical reasoning can benefit various technical applications. The model's resilience to diverse system prompts also makes it a strong candidate for chatbot development and role-playing scenarios.