willhx/Qwen3-8B-Base-SeaSFT-Search
The willhx/Qwen3-8B-Base-SeaSFT-Search model is an 8.2 billion parameter causal language model from the Qwen3 series, pre-trained on 36 trillion tokens across 119 languages. It features an expanded, higher-quality pre-training corpus with a rich mix of coding, STEM, reasoning, and multilingual data. This base model incorporates architectural refinements and a three-stage pre-training process, excelling in broad language modeling, reasoning skills, and long-context comprehension up to 32,768 tokens.
Loading preview...
Qwen3-8B-Base-SeaSFT-Search Overview
This model is an 8.2 billion parameter base causal language model from the Qwen3 series, developed by the Qwen Team. It builds upon significant advancements over previous Qwen models, focusing on an expanded and higher-quality pre-training corpus. The model was trained on an extensive 36 trillion tokens covering 119 languages, significantly tripling the language coverage of its predecessors and including a richer mix of high-quality data such as coding, STEM, reasoning, and multilingual content.
Key Capabilities and Features
- Expanded Pre-training Corpus: Trained on 36 trillion tokens across 119 languages, with a focus on high-quality data for coding, STEM, reasoning, and multilingual tasks.
- Architectural Refinements: Incorporates advanced training techniques and architectural improvements, including qk layernorm, for enhanced stability and performance.
- Three-Stage Pre-training: Utilizes a structured pre-training approach:
- Stage 1: Broad language modeling and general knowledge.
- Stage 2: Improved reasoning skills in STEM, coding, and logical reasoning.
- Stage 3: Enhanced long-context comprehension, supporting sequence lengths up to 32,768 tokens.
- Optimized Hyperparameter Tuning: Benefits from scaling law studies to systematically tune hyperparameters for better training dynamics and performance.
Model Specifications
- Parameters: 8.2 billion (6.95 billion non-embedding parameters)
- Context Length: 32,768 tokens
- Layers: 36
- Attention Heads (GQA): 32 for Q, 8 for KV
This model is designed for broad language understanding and generation tasks, with a strong foundation in reasoning and multilingual capabilities due to its extensive and diverse training data.