Ronaldodev/Qwen2.5-0.5B
Ronaldodev/Qwen2.5-0.5B is a 0.49 billion parameter causal language model from the Qwen2.5 series, developed by Qwen. This base model features a 32,768 token context length and is designed with transformers architecture including RoPE, SwiGLU, and RMSNorm. It offers significantly improved capabilities in coding, mathematics, instruction following, and long-text generation, with multilingual support for over 29 languages. It is intended for further post-training applications like SFT or RLHF rather than direct conversational use.
Loading preview...
Qwen2.5-0.5B Overview
Ronaldodev/Qwen2.5-0.5B is a base causal language model from the Qwen2.5 series, developed by Qwen. This model, with 0.49 billion parameters and a 32,768 token context length, is built upon a transformer architecture incorporating RoPE, SwiGLU, and RMSNorm. It represents an advancement over previous Qwen2 models, offering enhanced capabilities across several key areas.
Key Capabilities & Improvements
- Expanded Knowledge & Performance: Significantly improved in coding and mathematics due to specialized expert models.
- Instruction Following: Enhanced instruction following, long-text generation (over 8K tokens), and understanding of structured data like tables.
- Structured Output: Better generation of structured outputs, particularly JSON.
- System Prompt Resilience: More robust to diverse system prompts, aiding in role-play and chatbot condition-setting.
- Multilingual Support: Supports over 29 languages, including Chinese, English, French, Spanish, and more.
Intended Use
This 0.5B model is a base language model and is not recommended for direct conversational use. Instead, it is designed as a foundation for further post-training, such as Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), or continued pretraining, to adapt it for specific applications.