ArchiveStudio/Qwen3-14B-Base

TEXT GENERATIONConcurrent Unit Cost:1Model Size:14BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3-14B-Base is a 14.8 billion parameter causal language model developed by Qwen, part of the Qwen3 series. It is pre-trained on 36 trillion tokens across 119 languages, featuring an expanded, high-quality corpus including coding, STEM, and reasoning data. The model incorporates architectural refinements and a three-stage pre-training process to enhance broad language modeling, reasoning skills, and long-context comprehension up to 32,768 tokens. It is designed for foundational language understanding and generation tasks.

Loading preview...

Qwen3-14B-Base Overview

Qwen3-14B-Base is a 14.8 billion parameter causal language model from the Qwen3 series, developed by Qwen. This base model is pre-trained and features a context length of 32,768 tokens. It builds upon advancements in training data, model architecture, and optimization techniques, offering significant improvements over its predecessor, Qwen2.5.

Key Capabilities & Features

  • Expanded Pre-training Corpus: Trained on 36 trillion tokens across 119 languages, tripling the language coverage of Qwen2.5. The corpus includes a rich mix of high-quality data, such as coding, STEM, reasoning, and multilingual content.
  • Architectural Refinements: Incorporates training techniques like global-batch load balancing loss (for MoE models) and qk layernorm for all models, enhancing stability and performance.
  • Three-stage Pre-training:
    • Stage 1: Focuses on broad language modeling and general knowledge acquisition.
    • Stage 2: Improves reasoning skills, including STEM, coding, and logical reasoning.
    • Stage 3: Extends training sequence lengths up to 32k tokens for enhanced long-context comprehension.
  • Scaling Law Guided Hyperparameter Tuning: Critical hyperparameters were systematically tuned across the pre-training pipeline for optimal training dynamics and performance.

Good For

  • Foundational language understanding and generation tasks.
  • Applications requiring broad multilingual support.
  • Tasks benefiting from improved reasoning capabilities and long-context processing.