ArchiveStudio/Qwen3-14B-Base
Qwen3-14B-Base is a 14.8 billion parameter causal language model developed by Qwen, part of the Qwen3 series. It is pre-trained on 36 trillion tokens across 119 languages, featuring an expanded, high-quality corpus including coding, STEM, and reasoning data. The model incorporates architectural refinements and a three-stage pre-training process to enhance broad language modeling, reasoning skills, and long-context comprehension up to 32,768 tokens. It is designed for foundational language understanding and generation tasks.
Loading preview...
Qwen3-14B-Base Overview
Qwen3-14B-Base is a 14.8 billion parameter causal language model from the Qwen3 series, developed by Qwen. This base model is pre-trained and features a context length of 32,768 tokens. It builds upon advancements in training data, model architecture, and optimization techniques, offering significant improvements over its predecessor, Qwen2.5.
Key Capabilities & Features
- Expanded Pre-training Corpus: Trained on 36 trillion tokens across 119 languages, tripling the language coverage of Qwen2.5. The corpus includes a rich mix of high-quality data, such as coding, STEM, reasoning, and multilingual content.
- Architectural Refinements: Incorporates training techniques like global-batch load balancing loss (for MoE models) and qk layernorm for all models, enhancing stability and performance.
- Three-stage Pre-training:
- Stage 1: Focuses on broad language modeling and general knowledge acquisition.
- Stage 2: Improves reasoning skills, including STEM, coding, and logical reasoning.
- Stage 3: Extends training sequence lengths up to 32k tokens for enhanced long-context comprehension.
- Scaling Law Guided Hyperparameter Tuning: Critical hyperparameters were systematically tuned across the pre-training pipeline for optimal training dynamics and performance.
Good For
- Foundational language understanding and generation tasks.
- Applications requiring broad multilingual support.
- Tasks benefiting from improved reasoning capabilities and long-context processing.