Mojank/Qwen2.5-0.5B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Mojank/Qwen2.5-0.5B is a 0.49 billion parameter causal language model from the Qwen2.5 series, developed by Qwen. This base model features a transformer architecture with a 32,768-token context length, designed for pretraining. It offers significant improvements in coding, mathematics, instruction following, and long text generation, serving as a foundation for further fine-tuning.

Loading preview...

Qwen2.5-0.5B: An Enhanced Foundation Model

Qwen2.5-0.5B is a 0.49 billion parameter base causal language model, part of the latest Qwen2.5 series developed by Qwen. This model builds upon the Qwen2 architecture, incorporating specialized expert models to significantly enhance its capabilities in coding and mathematics. It features a transformer architecture with RoPE, SwiGLU, RMSNorm, and a substantial context length of 32,768 tokens.

Key Capabilities & Improvements

  • Enhanced Knowledge & Reasoning: Greatly improved performance in coding and mathematical tasks.
  • Instruction Following: Significant advancements in adhering to instructions and generating structured outputs like JSON.
  • Long Text Generation: Improved ability to generate texts over 8K tokens and understand structured data such as tables.
  • Multilingual Support: Offers support for over 29 languages, including Chinese, English, French, Spanish, and more.
  • Robustness: More resilient to diverse system prompts, aiding in role-play and chatbot condition-setting.

Usage Recommendations

This 0.5B model is a base language model intended for pretraining. It is not recommended for direct conversational use without further post-training steps such as Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), or continued pretraining. For more details, refer to the official blog and GitHub repository.