ayushrafukia24/Qwen2.5-0.5B

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen2.5-0.5B is a 0.49 billion parameter causal language model from the Qwen2.5 series, developed by the Qwen Team. This base model features a transformer architecture with RoPE, SwiGLU, RMSNorm, and a 32,768-token context length. It is designed for pretraining and serves as a foundation for further fine-tuning, offering enhanced capabilities in coding, mathematics, instruction following, and long-text generation compared to its predecessor, Qwen2.

Loading preview...

Qwen2.5-0.5B Overview

Qwen2.5-0.5B is a base causal language model, part of the latest Qwen2.5 series developed by the Qwen Team. This model, with 0.49 billion parameters and a 32,768-token context length, builds upon the Qwen2 architecture, incorporating improvements in several key areas. It is intended for pretraining and subsequent fine-tuning rather than direct conversational use.

Key Capabilities

  • Enhanced Knowledge & Reasoning: Significantly improved capabilities in coding and mathematics, leveraging specialized expert models.
  • Instruction Following: Better adherence to instructions and more resilient to diverse system prompts, aiding in role-play and chatbot condition-setting.
  • Long-Text Generation: Improved ability to generate texts exceeding 8,000 tokens.
  • Structured Data Understanding: Enhanced understanding of structured data, such as tables, and improved generation of structured outputs, including JSON.
  • Multilingual Support: Supports over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, and Arabic.

Good For

  • Foundation for Fine-tuning: Ideal for researchers and developers looking to apply post-training techniques like SFT, RLHF, or continued pretraining.
  • Specialized Applications: Suitable as a base for models requiring strong performance in coding, mathematical reasoning, or structured data processing.
  • Long-Context Tasks: Useful for applications that benefit from processing and generating long sequences of text up to 32,768 tokens.