Qwen/Qwen2-0.5B

Hugging Face
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 31, 2024License:apache-2.0Architecture:Transformer0.2K Open Weights Featherless Exclusive Warm

Qwen2-0.5B is a 0.5 billion parameter base language model from the Qwen2 series developed by Qwen. This Transformer-based model features SwiGLU activation, attention QKV bias, and group query attention, alongside an improved tokenizer for multiple natural languages and codes. It serves as a foundational model designed for further post-training applications like SFT, RLHF, or continued pretraining, rather than direct text generation.

Loading preview...

Qwen2-0.5B: A Compact Base Language Model

Qwen2-0.5B is a 0.5 billion parameter base model within the new Qwen2 series of large language models, developed by Qwen. This model is built upon the Transformer architecture, incorporating features such as SwiGLU activation, attention QKV bias, and group query attention. It also utilizes an enhanced tokenizer optimized for a wide array of natural languages and programming codes.

Key Capabilities & Design

  • Architecture: Decoder-only Transformer with SwiGLU activation and group query attention.
  • Tokenizer: Improved tokenizer designed for multilingual and multi-code adaptability.
  • Foundation Model: Intended as a base model for subsequent fine-tuning (e.g., SFT, RLHF) or continued pretraining.

Performance Insights

While Qwen2-0.5B is a base model, its performance is evaluated across diverse benchmarks including language understanding (MMLU, MMLU-Pro), coding (HumanEval, MBPP), mathematics (GSM8K, MATH), and multilingual tasks (C-Eval, CMMLU). In comparison to models like Phi-2 and Gemma-2B, Qwen2-0.5B demonstrates competitive, and in some cases, superior performance, particularly in Chinese language tasks like C-Eval and CMMLU, despite its smaller parameter count (0.35B non-embedding parameters).

Good for

  • Developing custom models: Ideal for researchers and developers looking to build specialized models through Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), or continued pretraining.
  • Exploring compact LLM capabilities: Suitable for understanding the performance of smaller, efficient language models on a variety of tasks.
  • Multilingual applications: Its improved tokenizer and evaluation on multilingual benchmarks suggest potential for applications requiring diverse language support after fine-tuning.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p