ArchiveStudio/Qwen3-0.6B-Base

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3-0.6B-Base is a 0.6 billion parameter causal language model from the Qwen3 series, developed by Qwen. Pre-trained on 36 trillion tokens across 119 languages, it features an expanded, high-quality corpus and architectural refinements like qk layernorm. This base model excels in broad language modeling, general knowledge acquisition, and long-context comprehension up to 32,768 tokens, making it suitable for diverse foundational NLP tasks.

Loading preview...

Qwen3-0.6B-Base Overview

Qwen3-0.6B-Base is a 0.6 billion parameter causal language model, part of the latest Qwen3 series. It builds upon significant advancements in training data, model architecture, and optimization techniques compared to its predecessor, Qwen2.5. This model is pre-trained on an extensive corpus of 36 trillion tokens covering 119 languages, with a rich mix of high-quality data including coding, STEM, reasoning, and multilingual content.

Key Features and Training

  • Expanded Pre-training Corpus: Utilizes 36 trillion tokens across 119 languages, significantly increasing language coverage and data quality.
  • Architectural Refinements: Incorporates advanced training techniques and architectural improvements, such as qk layernorm, for enhanced stability and performance.
  • Three-stage Pre-training: The model undergoes a structured pre-training process:
    • Stage 1: Focuses on broad language modeling and general knowledge.
    • Stage 2: Improves reasoning skills, including STEM, coding, and logical reasoning.
    • Stage 3: Enhances long-context comprehension, extending training sequence lengths up to 32,768 tokens.
  • Scaling Law Guided Hyperparameter Tuning: Critical hyperparameters are systematically tuned for optimal training dynamics and performance across different model scales.

Model Specifications

  • Parameters: 0.6 billion (0.44 billion non-embedding)
  • Layers: 28
  • Attention Heads (GQA): 16 for Q, 8 for KV
  • Context Length: 32,768 tokens

Use Cases

This base model is well-suited for foundational natural language processing tasks requiring broad language understanding, general knowledge, and the ability to process long contexts. Its comprehensive training makes it a strong candidate for further fine-tuning on specialized downstream applications.