cold368733/Qwen2.5-0.5B-Instruct
cold368733/Qwen2.5-0.5B-Instruct is a 0.49 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. It features a 32,768 token context length and is designed with improved capabilities in coding, mathematics, and instruction following. This model excels at generating long texts, understanding structured data like tables, and producing structured outputs such as JSON, while also offering robust multilingual support across 29 languages.
Loading preview...
Qwen2.5-0.5B-Instruct Overview
This model is the instruction-tuned 0.5 billion parameter variant from the Qwen2.5 series, developed by Qwen. It builds upon the Qwen2 architecture with significant enhancements across several key areas, making it a versatile choice for various NLP tasks. The model utilizes a transformer architecture with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings.
Key Capabilities & Improvements
- Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics due to specialized expert models.
- Instruction Following: Demonstrates substantial improvements in adhering to instructions and generating coherent, long-form text (up to 8K tokens).
- Structured Data Handling: Better at understanding structured data, including tables, and generating structured outputs like JSON.
- System Prompt Resilience: More robust against diverse system prompts, improving role-play and chatbot condition-setting.
- Long-Context Support: Features a full context length of 32,768 tokens, with the ability to generate up to 8,192 tokens.
- Multilingual Support: Supports over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, and Vietnamese.
Architecture Details
This specific model has 0.49 billion parameters (0.36B non-embedding), 24 layers, and 14 attention heads for Q with 2 for KV (GQA). For more detailed evaluation results and technical insights, refer to the official Qwen2.5 blog.