willhx/Qwen3-8B-Base-Math-SeaSFT-Search

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 21, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

willhx/Qwen3-8B-Base-Math-SeaSFT-Search is an 8.2 billion parameter causal language model from the Qwen3 series, pre-trained on 36 trillion tokens across 119 languages. This base model incorporates architectural refinements like qk layernorm and a three-stage pre-training process, focusing on broad language modeling, reasoning skills (STEM, coding, logical reasoning), and long-context comprehension up to 32,768 tokens. It is designed for foundational language understanding and generation tasks, serving as a robust base for further fine-tuning.

Loading preview...

Model Overview

willhx/Qwen3-8B-Base-Math-SeaSFT-Search is an 8.2 billion parameter causal language model, part of the latest Qwen3 series. This base model is pre-trained, meaning it provides a strong foundation for various natural language processing tasks before any specific instruction tuning. It features a substantial context length of 32,768 tokens, enabling it to process and generate longer sequences of text.

Key Enhancements in Qwen3

Qwen3 builds upon previous Qwen models with several significant improvements:

  • Expanded Pre-training Corpus: Trained on 36 trillion tokens covering 119 languages, with a rich mix of high-quality data including coding, STEM, reasoning, and multilingual content.
  • Architectural Refinements: Incorporates advanced training techniques and architectural changes, such as qk layernorm, to enhance stability and performance.
  • Three-stage Pre-training: A structured training approach that first focuses on general knowledge, then improves reasoning skills (STEM, coding, logical reasoning), and finally extends long-context comprehension.
  • Scaling Law Guided Tuning: Hyperparameters are systematically tuned across the pre-training pipeline for optimal performance at different model scales.

Use Cases

This base model is suitable for developers and researchers looking for a powerful, pre-trained language model to:

  • Develop custom applications requiring strong language understanding and generation capabilities.
  • Fine-tune for specific downstream tasks such as summarization, translation, or question answering.
  • Explore advanced reasoning and long-context processing in various domains.