Gaery0611/Qwen2.5-0.5B-Instruct
Gaery0611/Qwen2.5-0.5B-Instruct is a 0.49 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen Team. This model features a 32,768 token context length and is significantly improved in coding, mathematics, and instruction following compared to its predecessor. It excels at generating long texts, understanding structured data like JSON, and offers robust multilingual support across 29 languages.
Loading preview...
Qwen2.5-0.5B-Instruct Overview
This repository hosts the instruction-tuned 0.5 billion parameter model from the Qwen2.5 series, developed by the Qwen Team. Qwen2.5 models represent a significant advancement over Qwen2, incorporating specialized expert models to enhance capabilities in key areas.
Key Capabilities & Improvements
- Enhanced Knowledge & Performance: Demonstrates significantly improved capabilities in coding and mathematics due to specialized expert models.
- Instruction Following: Features substantial improvements in adhering to instructions and generating structured outputs, particularly JSON.
- Long Text Generation: Excels at generating extended texts, supporting outputs up to 8,192 tokens.
- Structured Data Understanding: Better at processing and understanding structured data, such as tables.
- Robustness: More resilient to diverse system prompts, improving role-play and chatbot condition-setting.
- Context Length: Supports a full context length of 32,768 tokens.
- Multilingual Support: Offers comprehensive support for over 29 languages, including Chinese, English, French, Spanish, and more.
Architecture & Features
This model is a causal language model built on transformers, incorporating RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings. It has 24 layers and 14 attention heads for Q with 2 for KV.
When to Use This Model
This model is particularly well-suited for applications requiring efficient instruction following, code generation, mathematical problem-solving, and structured output generation in a multilingual context, especially where a smaller, performant model is desired.