prompt-agnostic-language-models/sweep_Qwen-1B_ppcl_lr1e-05_js0.1
The prompt-agnostic-language-models/sweep_Qwen-1B_ppcl_lr1e-05_js0.1 model is a 0.8 billion parameter causal language model from the Qwen3 series, developed by Qwen Team. It is pre-trained on 36 trillion tokens across 119 languages, with a focus on high-quality data including coding, STEM, reasoning, and long-context comprehension up to 32,768 tokens. This model incorporates advanced training techniques like global-batch load balancing and qk layernorm, and utilizes a three-stage pre-training process to enhance broad language modeling, reasoning skills, and long-context understanding.
Loading preview...
Model Overview: Qwen3-0.6B-Base
This model is part of the Qwen3 series, a new generation of large language models developed by the Qwen Team. It is a pre-trained causal language model with 0.6 billion parameters (0.44 billion non-embedding parameters) and supports a substantial context length of 32,768 tokens. The Qwen3 series builds upon previous Qwen models with significant advancements in training data, architecture, and optimization.
Key Capabilities & Improvements:
- Expanded Pre-training Corpus: Trained on an extensive 36 trillion tokens across 119 languages, tripling the language coverage of its predecessor, Qwen2.5. The dataset includes a rich mix of high-quality data for coding, STEM, reasoning, and multilingual tasks.
- Advanced Training Techniques: Incorporates architectural refinements such as global-batch load balancing loss for MoE models and qk layernorm for all models, enhancing stability and overall performance.
- Three-stage Pre-training: A structured approach focusing on broad language modeling and general knowledge (Stage 1), improving reasoning skills like STEM and coding (Stage 2), and enhancing long-context comprehension by extending sequence lengths up to 32k tokens (Stage 3).
- Scaling Law Guided Hyperparameter Tuning: Critical hyperparameters are systematically tuned across the three-stage pipeline for optimal training dynamics and performance.
When to Use This Model:
This model is suitable for applications requiring a compact yet capable language model with strong multilingual support and enhanced reasoning abilities, particularly in scenarios benefiting from a large context window. Its pre-training focus on coding, STEM, and logical reasoning makes it a strong candidate for tasks in these domains.