andrewyang4ever/Qwen3-1.7B-Base

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 10, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3-1.7B-Base is a 1.7 billion parameter causal language model from the Qwen3 series, developed by Qwen. Pre-trained on 36 trillion tokens across 119 languages, it features an expanded, high-quality corpus and incorporates architectural refinements like qk layernorm. This model excels in broad language modeling, general knowledge acquisition, and reasoning skills, with a context length of 32,768 tokens.

Loading preview...

Qwen3-1.7B-Base Overview

Qwen3-1.7B-Base is a 1.7 billion parameter causal language model, part of the Qwen3 series developed by Qwen. This model builds upon advancements in training data, architecture, and optimization techniques, offering significant improvements over previous Qwen iterations.

Key Capabilities and Features

  • Expanded Pre-training Corpus: Trained on an extensive 36 trillion tokens across 119 languages, tripling the language coverage of Qwen2.5. The corpus includes a rich mix of high-quality data, such as coding, STEM, reasoning, books, multilingual content, and synthetic data.
  • Architectural Refinements: Incorporates advanced training techniques and architectural improvements, including qk layernorm for all models, enhancing stability and overall performance.
  • Three-stage Pre-training: The training process is structured in three stages:
    • Stage 1: Focuses on broad language modeling and general knowledge.
    • Stage 2: Improves reasoning skills, including STEM, coding, and logical reasoning.
    • Stage 3: Enhances long-context comprehension by extending training sequence lengths up to 32,768 tokens.
  • Scaling Law Guided Tuning: Hyperparameters are systematically tuned across the three-stage pipeline using comprehensive scaling law studies, leading to improved training dynamics and performance.

Model Specifications

  • Parameters: 1.7 billion (1.4 billion non-embedding parameters)
  • Layers: 28
  • Attention Heads (GQA): 16 for Q, 8 for KV
  • Context Length: 32,768 tokens

When to Use This Model

Qwen3-1.7B-Base is suitable for applications requiring strong general language understanding, reasoning, and multilingual capabilities within a compact parameter size. Its extended context length makes it effective for tasks involving longer inputs. For detailed evaluation results and further information, refer to the Qwen3 blog and GitHub repository.