Archivefusion/Qwen2-0.5B-Instruct
Archivefusion/Qwen2-0.5B-Instruct is a 0.5 billion parameter instruction-tuned causal language model from the Qwen2 series, developed by Qwen. Built on a Transformer architecture with SwiGLU activation and group query attention, it features an improved tokenizer for multilingual and code adaptability. This model is optimized for a broad range of tasks including language understanding, generation, coding, mathematics, and reasoning, demonstrating competitive performance against other open-source models.
Loading preview...
Qwen2-0.5B-Instruct Overview
Archivefusion/Qwen2-0.5B-Instruct is a 0.5 billion parameter instruction-tuned model from the new Qwen2 series, developed by Qwen. This series includes a range of base and instruction-tuned models, with Qwen2 generally outperforming many open-source models and showing competitiveness against proprietary models across various benchmarks.
Key Capabilities & Features
- Architecture: Based on the Transformer architecture, incorporating SwiGLU activation, attention QKV bias, and group query attention.
- Improved Tokenizer: Designed for enhanced adaptability across multiple natural languages and programming codes.
- Training: Pretrained on extensive datasets, followed by post-training using both supervised finetuning and direct preference optimization.
- Performance: Demonstrates significant improvements over its predecessor, Qwen1.5-0.5B-Chat, in key benchmarks:
- MMLU: Improved from 35.0 to 37.9
- HumanEval: Improved from 9.1 to 17.1
- GSM8K: Improved from 11.3 to 40.1
- C-Eval: Improved from 37.2 to 45.2
- IFEval (Prompt Strict-Acc.): Improved from 14.6 to 20.0
Use Cases
This model is suitable for applications requiring strong performance in:
- Language understanding and generation
- Multilingual tasks
- Coding assistance
- Mathematical problem-solving
- General reasoning tasks