ReliquaryForge/qwen3-4b-base-dapo-v4

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 18, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

ReliquaryForge/qwen3-4b-base-dapo-v4 is a 4.0 billion parameter causal language model from the Qwen3 series, pre-trained on 36 trillion tokens across 119 languages. Developed by Qwen, this base model incorporates architectural refinements and a three-stage pre-training process to enhance broad language modeling, reasoning skills, and long-context comprehension up to 32,768 tokens. It is designed for general language understanding and generation tasks, leveraging an expanded, high-quality multilingual corpus.

Loading preview...

Qwen3-4B-Base Overview

ReliquaryForge/qwen3-4b-base-dapo-v4 is a 4.0 billion parameter causal language model, part of the latest Qwen3 series developed by Qwen. This base model is pre-trained on an extensive corpus of 36 trillion tokens covering 119 languages, significantly expanding on previous Qwen versions with a richer mix of high-quality data including coding, STEM, reasoning, and synthetic content.

Key Features & Improvements

  • Expanded Pre-training Corpus: Utilizes 36 trillion tokens across 119 languages, tripling the language coverage and enhancing data quality for diverse tasks.
  • Architectural Refinements: Incorporates advanced training techniques and architectural improvements, such as qk layernorm, for improved stability and performance.
  • Three-stage Pre-training: Employs a structured pre-training approach:
    • Stage 1: Focuses on broad language modeling and general knowledge.
    • Stage 2: Enhances reasoning skills, including STEM, coding, and logical reasoning.
    • Stage 3: Improves long-context comprehension, extending training sequence lengths up to 32,768 tokens.
  • Context Length: Supports a substantial context window of 32,768 tokens.

Ideal Use Cases

  • General Language Understanding: Suitable for tasks requiring broad linguistic comprehension.
  • Multilingual Applications: Benefits from extensive multilingual pre-training for diverse language support.
  • Reasoning Tasks: Improved reasoning capabilities make it suitable for STEM, coding, and logical problem-solving.
  • Long-Context Processing: Its 32k context length is advantageous for applications requiring analysis or generation over extended texts.