cosmos1030/gmp-kd3e-1-s70pct-lr1e-4_20260827_230324

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 9, 2026Architecture:Transformer Featherless Exclusive Cold

The cosmos1030/gmp-kd3e-1-s70pct-lr1e-4_20260827_230324 model is an 8 billion parameter variant of the Qwen3 architecture, featuring a 32768 token context length. This model was developed by cosmos1030 and utilizes a TR-GMP capped-PGD 'during growth' variant with 70% unstructured data. It is specifically optimized for performance across benchmarks, achieving 5/5 on internal evaluations.

Loading preview...

Model Overview

The cosmos1030/gmp-kd3e-1-s70pct-lr1e-4_20260827_230324 is an 8 billion parameter language model based on the Qwen3 architecture, developed by cosmos1030. It incorporates a significant architectural modification, specifically a TR-GMP capped-PGD 'during growth' variant with a high proportion of 70% unstructured data in its training. This unique training methodology aims to enhance its overall performance.

Key Capabilities

  • Qwen3 Architecture: Leverages the robust foundation of the Qwen3 model family.
  • GMP Training Variant: Utilizes a specialized 'during growth' variant of TR-GMP capped-PGD, indicating a focus on efficient and effective model development.
  • High Unstructured Data Ratio: Trained with 70% unstructured data, potentially improving its ability to handle diverse and less-structured inputs.
  • Benchmark Performance: Achieved a perfect score of 5/5 on internal benchmarks, suggesting strong performance across evaluated metrics.
  • Context Length: Supports a substantial context window of 32768 tokens, enabling processing of longer inputs and maintaining coherence over extended conversations or documents.

Good For

  • Applications requiring a Qwen3-based model with specialized training for potentially improved efficiency or specific performance characteristics.
  • Use cases where handling a high proportion of unstructured data is beneficial.
  • Scenarios demanding a large context window for complex tasks or long-form content generation.