konizquants/Qwen3-8B-Base

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3-8B-Base is an 8.2 billion parameter causal language model developed by Qwen, pre-trained on 36 trillion tokens across 119 languages with a 32,768 token context length. This model incorporates advanced training techniques and architectural refinements, including qk layernorm and a three-stage pre-training process. It is designed for broad language modeling and general knowledge acquisition, with a focus on improving reasoning skills like STEM, coding, and logical reasoning.

Loading preview...

Overview

Qwen3-8B-Base is a pre-trained causal language model from the Qwen3 series, featuring 8.2 billion parameters and a 32,768 token context length. It builds upon the Qwen2.5 series with significant advancements in its training corpus and architectural design. The model was pre-trained on an expanded, higher-quality dataset of 36 trillion tokens covering 119 languages, tripling the language coverage of its predecessor. This corpus includes a rich mix of coding, STEM, reasoning, book, multilingual, and synthetic data.

Key Improvements and Training

Qwen3-8B-Base incorporates several training and architectural refinements, such as qk layernorm for improved stability and performance. Its development involved a three-stage pre-training process:

  • Stage 1: Focused on broad language modeling and general knowledge.
  • Stage 2: Enhanced reasoning skills, including STEM, coding, and logical reasoning.
  • Stage 3: Extended long-context comprehension by training with sequence lengths up to 32k tokens.

Hyperparameter tuning was guided by comprehensive scaling law studies across these stages, optimizing learning rate schedulers and batch sizes for better training dynamics. For detailed evaluation and performance metrics, refer to the Qwen3 blog.