ArchiveStudio/Qwen2.5-0.5B-Instruct

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 31, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ArchiveStudio/Qwen2.5-0.5B-Instruct is a 0.49 billion parameter instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. This model features a 32,768 token context length and is designed with improved capabilities in coding, mathematics, and instruction following. It excels at generating long texts, understanding structured data, and producing structured outputs like JSON, while also offering multilingual support for over 29 languages.

Loading preview...

Qwen2.5-0.5B-Instruct Overview

ArchiveStudio/Qwen2.5-0.5B-Instruct is an instruction-tuned model from the Qwen2.5 series, a significant advancement over previous Qwen models. This particular variant is a compact 0.49 billion parameter causal language model, built on a transformer architecture with a substantial 32,768 token context length and an 8,192 token generation capacity.

Key Capabilities & Improvements

  • Enhanced Knowledge & Reasoning: Demonstrates significantly improved capabilities in coding and mathematics, leveraging specialized expert models.
  • Instruction Following: Shows substantial improvements in adhering to instructions and understanding diverse system prompts, beneficial for role-play and chatbot implementations.
  • Structured Data Handling: Excels at understanding structured data, such as tables, and generating structured outputs, particularly JSON.
  • Long Text Generation: Capable of generating extended texts, supporting outputs over 8,000 tokens.
  • Multilingual Support: Provides robust support for over 29 languages, including major global languages like Chinese, English, French, Spanish, and Japanese.

Architecture & Training

The model utilizes a transformer architecture incorporating features like RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings. It has 24 layers and 14 attention heads for Q and 2 for KV (GQA). This instruction-tuned model has undergone both pretraining and post-training stages to optimize its performance across various tasks.