cosmos1030/gmp-kd3e-1-s50pct-lr5e-5_20260828_125050

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 9, 2026Architecture:Transformer Featherless Exclusive Cold

The cosmos1030/gmp-kd3e-1-s50pct-lr5e-5_20260828_125050 model is a 4 billion parameter language model based on the Qwen3-4B architecture. It was trained using 50% unstructured, uncapped PGD with a reverse-KL OPD objective, and features a 32768 token context length. This model is a result of specific experimental training configurations, making it distinct for research into advanced policy gradient methods.

Loading preview...

Model Overview

The cosmos1030/gmp-kd3e-1-s50pct-lr5e-5_20260828_125050 is a 4 billion parameter language model derived from the Qwen3-4B architecture. It was developed by cosmos1030 and trained with a specific experimental setup involving 50% unstructured, uncapped Policy Gradient Descent (PGD) combined with a reverse-KL On-Policy Distribution (OPD) objective. The training process completed 2048 steps, though evaluation encountered an issue after the first benchmark.

Key Characteristics

  • Architecture: Based on Qwen3-4B.
  • Parameter Count: 4 billion parameters.
  • Context Length: Supports a context window of 32768 tokens.
  • Training Methodology: Utilizes a unique training regimen with 50% unstructured, uncapped PGD and a reverse-KL OPD objective.
  • Experimental Focus: This model represents an archived experimental run, primarily serving as a record of a specific training configuration rather than a fully evaluated general-purpose model.

Intended Use Cases

This model is primarily suited for:

  • Research and Development: Ideal for researchers studying advanced policy gradient methods, particularly those involving unstructured PGD and reverse-KL OPD objectives.
  • Comparative Analysis: Can be used to compare the effects of its specific training methodology against other models or training paradigms.
  • Understanding Training Dynamics: Provides a snapshot of a completed experimental training run, useful for analyzing training logs and outcomes (though its evaluation data is limited).