ArchiveStudio/Qwen2.5-1.5B
ArchiveStudio/Qwen2.5-1.5B is a 1.54 billion parameter causal language model from the Qwen2.5 series, developed by Qwen. This transformer-based model features a 32,768 token context length and is significantly improved in knowledge, coding, and mathematics compared to its predecessors. It excels at instruction following, generating long texts, understanding structured data like JSON, and offers robust multilingual support for over 29 languages.
Loading preview...
Overview
ArchiveStudio/Qwen2.5-1.5B is a 1.54 billion parameter base causal language model from the Qwen2.5 series, developed by Qwen. This model is built on a transformer architecture, incorporating RoPE, SwiGLU, and RMSNorm, and supports a substantial context length of 32,768 tokens. It represents an advancement over the Qwen2 series, focusing on enhanced capabilities across several key areas.
Key Capabilities & Improvements
- Expanded Knowledge & Specialized Skills: Significantly improved in general knowledge, coding, and mathematics, benefiting from specialized expert models.
- Instruction Following & Generation: Demonstrates enhanced instruction following, improved generation of long texts (over 8K tokens), and better understanding and generation of structured data, particularly JSON.
- Robustness: More resilient to diverse system prompts, which aids in role-play implementation and setting chatbot conditions.
- Multilingual Support: Offers comprehensive support for over 29 languages, including major global languages like Chinese, English, French, Spanish, and Japanese.
- Long-Context Handling: While the base model supports 32,768 tokens, the Qwen2.5 series generally supports up to 128K tokens and can generate up to 8K tokens.
Important Note
This specific model is a base language model and is not recommended for direct conversational use. It is intended as a foundation for further post-training, such as Supervised Fine-Tuning (SFT) or Reinforcement Learning from Human Feedback (RLHF), to adapt it for specific applications.