prompt-agnostic-language-models/sweep_Qwen-1B_ppcl_lr1e-05_js100.0

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The prompt-agnostic-language-models/sweep_Qwen-1B_ppcl_lr1e-05_js100.0 is a 0.8 billion parameter causal language model from the Qwen3 series, developed by Qwen. This base model is pre-trained on an expanded 36 trillion token corpus covering 119 languages, incorporating architectural refinements like qk layernorm for improved stability. It is designed for broad language modeling and general knowledge acquisition, with a focus on enhancing reasoning skills and long-context comprehension up to 32,768 tokens through a three-stage pre-training process.

Loading preview...

Qwen3-0.6B-Base Model Overview

This model is part of the Qwen3 series, representing the latest generation of large language models from Qwen. It is a pre-trained causal language model with 0.6 billion parameters (0.44 billion non-embedding parameters) and a context length of 32,768 tokens. The Qwen3 series builds upon advancements in training data, model architecture, and optimization techniques, offering significant improvements over previous iterations.

Key Capabilities & Features

  • Expanded Pre-training Corpus: Trained on 36 trillion tokens across 119 languages, tripling the language coverage of Qwen2.5. The dataset includes a rich mix of high-quality data such as coding, STEM, reasoning, and multilingual content.
  • Architectural Refinements: Incorporates training techniques and architectural improvements like qk layernorm for enhanced stability and performance.
  • Three-stage Pre-training:
    • Stage 1: Focuses on broad language modeling and general knowledge.
    • Stage 2: Improves reasoning skills, including STEM, coding, and logical reasoning.
    • Stage 3: Enhances long-context comprehension by extending training sequence lengths.
  • Scaling Law Guided Hyperparameter Tuning: Utilizes comprehensive scaling law studies to systematically tune hyperparameters for optimal training dynamics and performance.

Good For

  • Applications requiring a compact yet capable causal language model.
  • Tasks benefiting from broad language understanding and general knowledge.
  • Scenarios where enhanced reasoning skills and long-context comprehension are valuable.
  • Developers looking for a model with a robust and diverse pre-training foundation.