ArchiveStudio/Qwen3-4B-Base

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3-4B-Base is a 4.0 billion parameter causal language model developed by Qwen, part of the Qwen3 series. This base model is pre-trained on an expanded 36 trillion token corpus covering 119 languages, with a rich mix of coding, STEM, and reasoning data. It incorporates architectural refinements and a three-stage pre-training approach to enhance general knowledge, reasoning skills, and long-context comprehension up to 32,768 tokens. It is designed for broad language modeling and foundational knowledge acquisition.

Loading preview...

Qwen3-4B-Base Overview

Qwen3-4B-Base is a 4.0 billion parameter causal language model from the Qwen3 series, developed by Qwen. This model represents the latest generation of Qwen's large language models, building on advancements in training data, architecture, and optimization techniques. It is a pre-trained base model, not instruction-tuned.

Key Characteristics & Improvements

  • Expanded Pre-training Corpus: Trained on an extensive 36 trillion tokens across 119 languages, significantly tripling the language coverage compared to Qwen2.5. The dataset includes a richer mix of high-quality data, such as coding, STEM, reasoning, and multilingual content.
  • Architectural Refinements: Incorporates advanced training techniques and architectural improvements, including qk layernorm, to enhance stability and overall performance.
  • Three-stage Pre-training: Utilizes a structured pre-training approach:
    • Stage 1: Focuses on broad language modeling and general knowledge acquisition.
    • Stage 2: Improves specialized reasoning skills, including STEM, coding, and logical reasoning.
    • Stage 3: Enhances long-context comprehension by extending training sequence lengths up to 32,768 tokens.
  • Scaling Law Guided Tuning: Critical hyperparameters were systematically tuned based on comprehensive scaling law studies across the three pre-training stages, optimizing training dynamics and performance.

Model Specifications

  • Type: Causal Language Model
  • Training Stage: Pretraining
  • Parameters: 4.0 billion (3.6 billion non-embedding)
  • Layers: 36
  • Attention Heads (GQA): 32 for Q, 8 for KV
  • Context Length: 32,768 tokens

For detailed evaluation results and further information, refer to the official Qwen3 blog and GitHub repository.