prompt-agnostic-language-models/sweep_Qwen-1B_ppcl_lr1e-05_js10.0

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The prompt-agnostic-language-models/sweep_Qwen-1B_ppcl_lr1e-05_js10.0 model is a 0.8 billion parameter causal language model from the Qwen3 series, developed by Qwen. It is pre-trained on 36 trillion tokens across 119 languages, incorporating advanced training techniques and architectural refinements like qk layernorm. This model is designed for broad language modeling, general knowledge acquisition, and improved reasoning skills, with a notable context length of 32,768 tokens.

Loading preview...

Qwen3-0.6B-Base Overview

This model is part of the Qwen3 series, a new generation of large language models from Qwen, building on advancements in training data, architecture, and optimization. While the specific model name sweep_Qwen-1B_ppcl_lr1e-05_js10.0 is a variant, the underlying architecture and training principles are derived from the Qwen3-0.6B-Base.

Key Improvements and Features

  • Expanded Pre-training Corpus: Trained on 36 trillion tokens across 119 languages, significantly increasing language coverage and data quality, including coding, STEM, reasoning, and multilingual data.
  • Advanced Training Techniques: Incorporates architectural refinements such as qk layernorm for improved stability and performance across all models.
  • Three-stage Pre-training: A structured approach focusing on:
    • Stage 1: Broad language modeling and general knowledge.
    • Stage 2: Enhanced reasoning skills (STEM, coding, logical reasoning).
    • Stage 3: Long-context comprehension, extending training sequence lengths up to 32,768 tokens.
  • Scaling Law Guided Hyperparameter Tuning: Critical hyperparameters are systematically tuned for better training dynamics and performance.

Model Specifications (Qwen3-0.6B-Base)

  • Type: Causal Language Model
  • Training Stage: Pretraining
  • Number of Parameters: 0.6B (0.44B non-embedding)
  • Context Length: 32,768 tokens

Use Cases

This model is suitable for applications requiring robust language understanding, generation, and reasoning across a wide range of languages, particularly benefiting from its extended context window and improved stability. Its pre-training on diverse high-quality data makes it versatile for general-purpose language tasks.