ErtasAI/Qwen2.5-1.5B-Instruct
ErtasAI/Qwen2.5-1.5B-Instruct is a 1.54 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. This model features a transformer architecture with RoPE, SwiGLU, and RMSNorm, supporting a full context length of 32,768 tokens and generating up to 8,192 tokens. It significantly improves capabilities in coding, mathematics, instruction following, long text generation, structured data understanding, and multilingual support across 29 languages.
Loading preview...
Qwen2.5-1.5B-Instruct Overview
This model is an instruction-tuned variant of the Qwen2.5 series, developed by Qwen, featuring 1.54 billion parameters. It builds upon the Qwen2 architecture with significant enhancements across several key areas, making it a versatile tool for various natural language processing tasks.
Key Capabilities and Improvements
- Enhanced Knowledge & Reasoning: Demonstrates significantly more knowledge and improved capabilities in coding and mathematics, leveraging specialized expert models.
- Instruction Following: Offers substantial improvements in adhering to instructions and generating structured outputs, including JSON.
- Long Context & Generation: Supports a full context length of 32,768 tokens and can generate responses up to 8,192 tokens.
- Multilingual Support: Provides robust support for over 29 languages, including major global languages like Chinese, English, French, Spanish, German, and Japanese.
- Structured Data Understanding: Better at understanding structured data, such as tables, and more resilient to diverse system prompts for improved role-play and chatbot implementation.
Architecture and Features
The model utilizes a transformer architecture incorporating RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings. It consists of 28 layers and 12 attention heads (GQA) for Q and 2 for KV. This specific repository hosts the instruction-tuned 1.5B model, ready for deployment and fine-tuning on custom datasets via platforms like Ertas AI.