ArchiveStudio/Qwen2.5-0.5B
ArchiveStudio/Qwen2.5-0.5B is a 0.49 billion parameter causal language model from the Qwen2.5 series, developed by Qwen. This base model features a transformer architecture with RoPE, SwiGLU, and RMSNorm, supporting a context length of 32,768 tokens. It serves as a foundation for further fine-tuning, offering enhanced capabilities in coding, mathematics, and multilingual support across 29 languages.
Loading preview...
Qwen2.5-0.5B Overview
ArchiveStudio/Qwen2.5-0.5B is a base causal language model from the Qwen2.5 series, featuring 0.49 billion parameters and a 32,768-token context length. Developed by Qwen, this model is built on a transformer architecture incorporating RoPE, SwiGLU, and RMSNorm. It represents an advancement over previous Qwen models, particularly in its foundational capabilities.
Key Capabilities & Improvements
- Enhanced Knowledge: Significantly improved understanding in coding and mathematics, leveraging specialized expert models.
- Instruction Following: Better at adhering to instructions and generating long texts (over 8K tokens).
- Structured Data Handling: Improved ability to understand and generate structured data, including JSON outputs.
- Multilingual Support: Supports over 29 languages, including Chinese, English, French, Spanish, and more.
- Robustness: More resilient to diverse system prompts, aiding in role-play and chatbot implementations.
Intended Use
This 0.5B model is a base language model and is not recommended for direct conversational use. It is designed as a strong foundation for further post-training applications such as Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), or continued pretraining. Developers can leverage its enhanced core capabilities to build specialized models tailored to specific tasks.