ArchiveStudio/Qwen2.5-3B
ArchiveStudio/Qwen2.5-3B is a 3.09 billion parameter causal language model from the Qwen2.5 series, developed by the Qwen Team. This base model features a transformer architecture with a 32,768-token context length. It offers significantly improved knowledge, coding, and mathematics capabilities compared to its predecessor, Qwen2, and supports over 29 languages. It is designed for pretraining and further fine-tuning for specific applications.
Loading preview...
Qwen2.5-3B: An Enhanced Base Language Model
ArchiveStudio/Qwen2.5-3B is a 3.09 billion parameter base causal language model, part of the latest Qwen2.5 series developed by the Qwen Team. This model builds upon the Qwen2 architecture, incorporating significant advancements across several key areas. It is designed for pretraining and serves as a foundation for further fine-tuning, such as SFT or RLHF, rather than direct conversational use.
Key Enhancements and Features
- Expanded Knowledge & Capabilities: Qwen2.5-3B demonstrates significantly more knowledge and improved performance in specialized domains like coding and mathematics, benefiting from expert model integration.
- Instruction Following & Text Generation: It shows substantial improvements in instruction following, generating long texts (up to 8K tokens), and understanding/generating structured data, including JSON outputs. The model is also more resilient to diverse system prompts, enhancing role-play and chatbot condition-setting.
- Long-Context Support: The model supports a full context length of 32,768 tokens and can generate outputs up to 8K tokens.
- Multilingual Support: Qwen2.5-3B offers robust multilingual capabilities, supporting over 29 languages, including major global languages like Chinese, English, French, Spanish, German, Japanese, and Korean.
- Architecture: It utilizes a transformer architecture with RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings.
Intended Use
This model is a base language model, primarily intended for developers to apply post-training techniques like supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or continued pretraining. It is not recommended for direct conversational use without further fine-tuning.