cosmos1030/gmp-kd3e-1-s70pct-lr1e-4_20260828_073925

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 9, 2026Architecture:Transformer Featherless Exclusive Cold

The cosmos1030/gmp-kd3e-1-s70pct-lr1e-4_20260828_073925 model is a 4 billion parameter Qwen3-based language model with a 32768 token context length. It was trained with 70% unstructured data and utilizes a capped PGD with reverse-KL OPD objective. This model demonstrates strong performance, achieving 5/5 on its benchmarks, making it suitable for tasks requiring robust language understanding and generation.

Loading preview...

Model Overview

The cosmos1030/gmp-kd3e-1-s70pct-lr1e-4_20260828_073925 is a 4 billion parameter language model built upon the Qwen3 architecture. It features a substantial context length of 32768 tokens, enabling it to process and generate longer sequences of text.

Training Methodology

This model was trained with a specific focus on data distribution, incorporating 70% unstructured data. The training process also employed a capped PGD (Proximal Gradient Descent) with a reverse-KL OPD (Optimal Policy Distribution) objective. This specialized training approach aims to enhance its performance and generalization capabilities.

Performance Highlights

The model has demonstrated strong results across its evaluation metrics, achieving a perfect 5/5 on its internal benchmarks. This indicates its effectiveness in the tasks it was designed for.

Key Characteristics

  • Architecture: Qwen3-based
  • Parameter Count: 4 billion
  • Context Length: 32768 tokens
  • Training Data: 70% unstructured data
  • Optimization: Capped PGD with reverse-KL OPD objective
  • Benchmark Performance: 5/5 on internal benchmarks

Intended Use Cases

Given its robust training and benchmark performance, this model is well-suited for applications requiring a capable language model with a focus on understanding and generating text from diverse, unstructured sources. Its large context window makes it particularly effective for tasks involving extensive document analysis or long-form content generation.