ErtasAI/Qwen2.5-0.5B-Instruct
ErtasAI/Qwen2.5-0.5B-Instruct is a 0.49 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. It features a 32,768 token context length and is optimized for enhanced knowledge, coding, mathematics, and instruction following. This model excels at generating long texts, understanding structured data like tables, and producing structured outputs such as JSON, with robust multilingual support across 29 languages.
Loading preview...
ErtasAI/Qwen2.5-0.5B-Instruct: Key Highlights
This model is an instruction-tuned variant of the Qwen2.5 series, developed by Qwen, featuring 0.49 billion parameters and a substantial 32,768 token context length. It builds upon the Qwen2 architecture with significant enhancements across several domains.
Key Capabilities & Improvements
- Expanded Knowledge & Specialized Skills: Demonstrates significantly more knowledge and greatly improved capabilities in coding and mathematics, leveraging specialized expert models.
- Enhanced Instruction Following: Offers substantial improvements in adhering to instructions, generating long texts (over 8K tokens), and understanding structured data (e.g., tables).
- Robust Structured Output: Excels at generating structured outputs, particularly JSON, and is more resilient to diverse system prompts, aiding in role-play and chatbot condition-setting.
- Long-Context Support: Supports a full context length of 32,768 tokens and can generate up to 8,192 tokens.
- Multilingual Proficiency: Provides strong multilingual support for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, and Arabic.
Architecture & Training
This model is a causal language model trained through pretraining and post-training stages. Its architecture includes transformers with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings. It comprises 24 layers and 14 attention heads (GQA) for Q and 2 for KV.
Use Cases
This model is particularly well-suited for applications requiring precise instruction following, structured data processing, code generation, mathematical problem-solving, and multilingual text generation, especially where a smaller, efficient model with a deep context window is beneficial.