prompt-agnostic-language-models/Qwen-1B_single_longer

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 9, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen-1B_single_longer is a 0.8 billion parameter causal language model from the Qwen3 series, developed by Qwen Team. This model is pre-trained on 36 trillion tokens across 119 languages, featuring an expanded, higher-quality corpus with a rich mix of coding, STEM, reasoning, and multilingual data. It incorporates architectural refinements and a three-stage pre-training process to enhance broad language modeling, reasoning skills, and long-context comprehension up to 32,768 tokens. Optimized for general language tasks, it offers improved stability and performance across various scales.

Loading preview...

Qwen3-0.6B-Base Overview

Qwen3-0.6B-Base is a 0.6 billion parameter causal language model from the latest Qwen3 series, developed by the Qwen Team. This model builds upon advancements in training data, architecture, and optimization techniques, offering significant improvements over previous Qwen iterations.

Key Capabilities & Features

  • Expanded Pre-training Corpus: Trained on an extensive 36 trillion tokens across 119 languages, tripling the language coverage of Qwen2.5. The dataset includes a richer mix of high-quality data, such as coding, STEM, reasoning, book, multilingual, and synthetic content.
  • Architectural Refinements: Incorporates training techniques and architectural improvements like qk layernorm for enhanced stability and performance.
  • Three-stage Pre-training: Utilizes a staged approach:
    • Stage 1: Focuses on broad language modeling and general knowledge acquisition.
    • Stage 2: Improves reasoning skills, including STEM, coding, and logical reasoning.
    • Stage 3: Enhances long-context comprehension by extending training sequence lengths up to 32,768 tokens.
  • Scaling Law Guided Tuning: Critical hyperparameters were systematically tuned across the pre-training pipeline for optimal training dynamics and performance.

Model Specifications

  • Type: Causal Language Model
  • Training Stage: Pretraining
  • Parameters: 0.6 Billion (0.44B non-embedding)
  • Context Length: 32,768 tokens

When to Use This Model

This model is suitable for applications requiring a compact yet capable language model with strong multilingual support and enhanced reasoning abilities, particularly for tasks benefiting from a long context window. Its comprehensive pre-training makes it versatile for general language understanding, code-related tasks, and logical reasoning.