abdullahshady/AqlPrime
The abdullahshady/AqlPrime model is an instruction-tuned 1.54 billion parameter causal language model from the Qwen2.5 series, developed by Qwen Team. It features a 32,768 token context length and is designed with transformers architecture including RoPE, SwiGLU, and RMSNorm. This model offers significant improvements in coding, mathematics, instruction following, long text generation, structured data understanding, and multilingual support across 29 languages.
Loading preview...
Model Overview
The abdullahshady/AqlPrime model is an instruction-tuned variant of the Qwen2.5 series, developed by the Qwen Team. This specific model has 1.54 billion parameters and supports a substantial context length of 32,768 tokens, with a generation capacity of up to 8,192 tokens. It is built on a transformer architecture incorporating RoPE, SwiGLU, RMSNorm, and attention QKV bias.
Key Capabilities & Improvements
Qwen2.5 models, including this 1.5B instruction-tuned version, offer several enhancements over previous iterations:
- Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics, leveraging specialized expert models.
- Instruction Following: Demonstrates better adherence to instructions and is more resilient to diverse system prompts, aiding in role-play and chatbot implementations.
- Long Text Generation: Excels at generating extended texts, supporting outputs over 8,000 tokens.
- Structured Data & Output: Improved understanding of structured data (e.g., tables) and generation of structured outputs, particularly JSON.
- Multilingual Support: Provides robust support for over 29 languages, including major global languages like Chinese, English, French, Spanish, and more.
Architecture Details
This model features 28 layers, 12 attention heads for queries, and 2 for key/value pairs (GQA). The non-embedding parameter count is 1.31 billion. For more in-depth information, users can refer to the official Qwen2.5 blog and GitHub repository.