prompt-agnostic-language-models/sweep_Qwen-1B_ppcl_lr1e-05_js0.1

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The prompt-agnostic-language-models/sweep_Qwen-1B_ppcl_lr1e-05_js0.1 model is a 0.8 billion parameter causal language model from the Qwen3 series, developed by Qwen Team. It is pre-trained on 36 trillion tokens across 119 languages, with a focus on high-quality data including coding, STEM, reasoning, and long-context comprehension up to 32,768 tokens. This model incorporates advanced training techniques like global-batch load balancing and qk layernorm, and utilizes a three-stage pre-training process to enhance broad language modeling, reasoning skills, and long-context understanding.

Loading preview...

Model Overview: Qwen3-0.6B-Base

This model is part of the Qwen3 series, a new generation of large language models developed by the Qwen Team. It is a pre-trained causal language model with 0.6 billion parameters (0.44 billion non-embedding parameters) and supports a substantial context length of 32,768 tokens. The Qwen3 series builds upon previous Qwen models with significant advancements in training data, architecture, and optimization.

Key Capabilities & Improvements:

  • Expanded Pre-training Corpus: Trained on an extensive 36 trillion tokens across 119 languages, tripling the language coverage of its predecessor, Qwen2.5. The dataset includes a rich mix of high-quality data for coding, STEM, reasoning, and multilingual tasks.
  • Advanced Training Techniques: Incorporates architectural refinements such as global-batch load balancing loss for MoE models and qk layernorm for all models, enhancing stability and overall performance.
  • Three-stage Pre-training: A structured approach focusing on broad language modeling and general knowledge (Stage 1), improving reasoning skills like STEM and coding (Stage 2), and enhancing long-context comprehension by extending sequence lengths up to 32k tokens (Stage 3).
  • Scaling Law Guided Hyperparameter Tuning: Critical hyperparameters are systematically tuned across the three-stage pipeline for optimal training dynamics and performance.

When to Use This Model:

This model is suitable for applications requiring a compact yet capable language model with strong multilingual support and enhanced reasoning abilities, particularly in scenarios benefiting from a large context window. Its pre-training focus on coding, STEM, and logical reasoning makes it a strong candidate for tasks in these domains.