tayaee/Qwen2.5-1.5B-Instruct-ko-Reasoning-alpha-smoke
The tayaee/Qwen2.5-1.5B-Instruct-ko-Reasoning-alpha-smoke model is an instruction-tuned 1.54 billion parameter causal language model from the Qwen2.5 series, developed by Qwen. It features a 32,768 token context length and is specifically enhanced for reasoning, coding, and mathematics. This model excels at instruction following, generating long texts, understanding structured data like JSON, and offers robust multilingual support across 29 languages.
Loading preview...
Overview
This model, tayaee/Qwen2.5-1.5B-Instruct-ko-Reasoning-alpha-smoke, is an instruction-tuned variant of the Qwen2.5 series, developed by Qwen. It is a 1.54 billion parameter causal language model built on a transformer architecture, featuring RoPE, SwiGLU, RMSNorm, and attention QKV bias. It supports a substantial context length of 32,768 tokens for input and can generate up to 8,192 tokens.
Key Capabilities & Improvements
Qwen2.5 models, including this 1.5B instruction-tuned version, offer significant advancements over previous Qwen2 iterations:
- Enhanced Reasoning & Technical Skills: Greatly improved capabilities in coding and mathematics due to specialized expert models.
- Superior Instruction Following: Demonstrates significant improvements in adhering to instructions and generating diverse outputs.
- Long Text Generation: Excels at producing long texts, supporting outputs over 8,000 tokens.
- Structured Data Handling: Better at understanding structured data (e.g., tables) and generating structured outputs, particularly JSON.
- Robustness: More resilient to varied system prompts, improving role-play and chatbot condition-setting.
- Multilingual Support: Provides comprehensive support for over 29 languages, including Korean, Chinese, English, Japanese, and many European languages.
When to Use This Model
This model is particularly well-suited for applications requiring strong instruction following, complex reasoning, code generation, and mathematical problem-solving. Its ability to handle long contexts and generate structured outputs makes it valuable for tasks like content creation, data extraction, and building sophisticated multilingual chatbots.