JMullings/Qwen2.5-0.5B-Instruct
JMullings/Qwen2.5-0.5B-Instruct is a 0.49 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. This model significantly improves capabilities in coding, mathematics, instruction following, and generating structured outputs like JSON. It supports a full context length of 32,768 tokens and is multilingual, covering over 29 languages.
Loading preview...
Qwen2.5-0.5B-Instruct Overview
This model is the instruction-tuned 0.5 billion parameter variant from the Qwen2.5 series, developed by Qwen. It builds upon the Qwen2 architecture, featuring transformers with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings. The model has 24 layers and a full context length of 32,768 tokens, with a generation capacity of 8,192 tokens.
Key Capabilities
- Enhanced Knowledge & Reasoning: Significantly improved performance in coding and mathematics, leveraging specialized expert models.
- Instruction Following: Demonstrates substantial improvements in adhering to instructions and generating long texts (over 8K tokens).
- Structured Data & Output: Excels at understanding structured data, such as tables, and generating structured outputs, particularly JSON.
- Robust System Prompt Handling: More resilient to diverse system prompts, which enhances role-play implementations and chatbot condition-setting.
- Multilingual Support: Provides support for over 29 languages, including Chinese, English, French, Spanish, German, and Japanese.
Good for
- Applications requiring strong instruction following and structured output generation, especially JSON.
- Tasks involving coding and mathematical reasoning, even at its compact size.
- Multilingual conversational agents or text generation in supported languages.
- Scenarios needing a model with a substantial context window (32K tokens) for understanding longer inputs.