ArchiveStudio/Qwen2.5-14B
ArchiveStudio/Qwen2.5-14B is a 14.7 billion parameter causal language model from the Qwen2.5 series, developed by the Qwen Team. This base model features a transformer architecture with RoPE, SwiGLU, and RMSNorm, supporting a context length of up to 131,072 tokens. It offers significantly improved capabilities in coding, mathematics, instruction following, and generating long, structured texts, making it suitable for further fine-tuning for specialized applications.
Loading preview...
Qwen2.5-14B Overview
ArchiveStudio/Qwen2.5-14B is a 14.7 billion parameter base causal language model from the Qwen2.5 series, developed by the Qwen Team. This model builds upon the Qwen2 architecture, incorporating transformer components like RoPE, SwiGLU, and RMSNorm. It is designed for pretraining and is not recommended for direct conversational use without further fine-tuning.
Key Capabilities & Improvements
- Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics, leveraging specialized expert models.
- Instruction Following: Demonstrates substantial improvements in adhering to instructions and generating structured outputs, including JSON.
- Long Text Generation: Better performance in generating texts over 8,000 tokens and understanding structured data like tables.
- Robustness: More resilient to diverse system prompts, which benefits role-play and chatbot condition-setting.
- Extended Context: Supports a long context length of up to 131,072 tokens and can generate up to 8,000 tokens.
- Multilingual Support: Offers support for over 29 languages, including major global languages.
Technical Specifications
- Parameters: 14.7 billion (13.1 billion non-embedding)
- Layers: 48
- Attention Heads (GQA): 40 for Q, 8 for KV
- Context Length: 131,072 tokens
This base model is intended for developers to apply post-training techniques such as SFT, RLHF, or continued pretraining to adapt it for specific use cases. More detailed evaluation results and performance benchmarks are available in the official blog and documentation.