ReliquaryForge/qwen3-4b-base-dapo-v4
ReliquaryForge/qwen3-4b-base-dapo-v4 is a 4.0 billion parameter causal language model from the Qwen3 series, pre-trained on 36 trillion tokens across 119 languages. Developed by Qwen, this base model incorporates architectural refinements and a three-stage pre-training process to enhance broad language modeling, reasoning skills, and long-context comprehension up to 32,768 tokens. It is designed for general language understanding and generation tasks, leveraging an expanded, high-quality multilingual corpus.
Loading preview...
Qwen3-4B-Base Overview
ReliquaryForge/qwen3-4b-base-dapo-v4 is a 4.0 billion parameter causal language model, part of the latest Qwen3 series developed by Qwen. This base model is pre-trained on an extensive corpus of 36 trillion tokens covering 119 languages, significantly expanding on previous Qwen versions with a richer mix of high-quality data including coding, STEM, reasoning, and synthetic content.
Key Features & Improvements
- Expanded Pre-training Corpus: Utilizes 36 trillion tokens across 119 languages, tripling the language coverage and enhancing data quality for diverse tasks.
- Architectural Refinements: Incorporates advanced training techniques and architectural improvements, such as qk layernorm, for improved stability and performance.
- Three-stage Pre-training: Employs a structured pre-training approach:
- Stage 1: Focuses on broad language modeling and general knowledge.
- Stage 2: Enhances reasoning skills, including STEM, coding, and logical reasoning.
- Stage 3: Improves long-context comprehension, extending training sequence lengths up to 32,768 tokens.
- Context Length: Supports a substantial context window of 32,768 tokens.
Ideal Use Cases
- General Language Understanding: Suitable for tasks requiring broad linguistic comprehension.
- Multilingual Applications: Benefits from extensive multilingual pre-training for diverse language support.
- Reasoning Tasks: Improved reasoning capabilities make it suitable for STEM, coding, and logical problem-solving.
- Long-Context Processing: Its 32k context length is advantageous for applications requiring analysis or generation over extended texts.