willhx/Qwen3-8B-Base-Math-SeaSFT-Search-TauSFT-Tau

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 21, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3-8B-Base-Math-SeaSFT-Search-TauSFT-Tau is an 8.2 billion parameter causal language model from the Qwen3 series, developed by the Qwen Team. Pre-trained on 36 trillion tokens across 119 languages, it features an expanded, high-quality corpus with a rich mix of coding, STEM, reasoning, and multilingual data. This base model incorporates architectural refinements and a three-stage pre-training process, focusing on broad language modeling, improved reasoning skills, and enhanced long-context comprehension up to 32,768 tokens.

Loading preview...

Model Overview

willhx/Qwen3-8B-Base-Math-SeaSFT-Search-TauSFT-Tau is an 8.2 billion parameter base causal language model from the Qwen3 series, developed by the Qwen Team. It builds upon the Qwen2.5 series with significant advancements in training data, architecture, and optimization techniques. This model is designed for broad language modeling and general knowledge acquisition, with a strong emphasis on reasoning and long-context understanding.

Key Highlights & Capabilities

  • Expanded Pre-training Corpus: Trained on an extensive 36 trillion tokens across 119 languages, tripling the language coverage of its predecessor. The dataset includes a rich mix of high-quality data, specifically focusing on coding, STEM, reasoning, books, multilingual content, and synthetic data.
  • Architectural Refinements: Incorporates advanced training techniques and architectural improvements, such as global-batch load balancing loss for MoE models and qk layernorm for all models, enhancing stability and overall performance.
  • Three-stage Pre-training: The training process is structured in three distinct stages:
    • Stage 1: Focuses on broad language modeling and general knowledge.
    • Stage 2: Improves reasoning skills, including STEM, coding, and logical reasoning.
    • Stage 3: Enhances long-context comprehension by extending training sequence lengths up to 32,768 tokens.
  • Scaling Law Guided Tuning: Critical hyperparameters were systematically tuned across the three-stage pipeline using comprehensive scaling law studies, optimizing training dynamics and performance.

Model Specifications

  • Type: Causal Language Model
  • Parameters: 8.2 billion (6.95 billion non-embedding)
  • Context Length: 32,768 tokens
  • Layers: 36
  • Attention Heads (GQA): 32 for Q, 8 for KV

When to Use This Model

This model is suitable for applications requiring strong general language understanding, advanced reasoning capabilities, and the ability to process long contexts. Its extensive multilingual training and focus on STEM and coding data make it a robust choice for diverse technical and linguistic tasks. For detailed evaluation results and further information, refer to the official Qwen3 blog and GitHub repository.