ArchiveStudio/Qwen3-8B-Base

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3-8B-Base is an 8.2 billion parameter causal language model from the Qwen3 series, developed by Qwen. Pre-trained on 36 trillion tokens across 119 languages, it features an expanded, high-quality corpus and architectural refinements like qk layernorm. This base model excels in broad language modeling, general knowledge acquisition, and improved reasoning skills, with a 32,768 token context length.

Loading preview...

Qwen3-8B-Base: A Foundation Model from the Qwen3 Series

Qwen3-8B-Base is an 8.2 billion parameter pre-trained causal language model, part of the latest generation of Qwen large language models. It builds upon significant advancements in training data, model architecture, and optimization techniques compared to its predecessor, Qwen2.5.

Key Capabilities and Improvements

  • Expanded Pre-training Corpus: Trained on an extensive 36 trillion tokens across 119 languages, tripling the language coverage of Qwen2.5. The dataset includes a rich mix of high-quality data, such as coding, STEM, reasoning, book, multilingual, and synthetic data.
  • Architectural Refinements: Incorporates advanced training techniques and architectural improvements, including qk layernorm, enhancing stability and overall performance.
  • Three-stage Pre-training: Utilizes a structured pre-training approach:
    • Stage 1: Focuses on broad language modeling and general knowledge.
    • Stage 2: Improves reasoning skills, including STEM, coding, and logical reasoning.
    • Stage 3: Enhances long-context comprehension by extending training sequence lengths up to 32,768 tokens.
  • Optimized Hyperparameter Tuning: Benefits from comprehensive scaling law studies to systematically tune critical hyperparameters for better training dynamics.

Model Specifications

  • Parameters: 8.2 billion (6.95 billion non-embedding parameters)
  • Context Length: 32,768 tokens
  • Layers: 36
  • Attention Heads (GQA): 32 for Q, 8 for KV

This model serves as a robust foundation for various natural language processing tasks, offering strong general knowledge and reasoning capabilities due to its extensive and diverse training.