cosmos1030/gmp-kd3e-1-s50pct-lr5e-5_20260828_125227
TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 9, 2026Architecture:Transformer Featherless Exclusive Cold
The cosmos1030/gmp-kd3e-1-s50pct-lr5e-5_20260828_125227 is a 4 billion parameter language model based on the Qwen3 architecture. It was trained with 50% unstructured data and a capped PGD with reverse-KL OPD objective. This model is notable for its specific training methodology aimed at exploring advanced optimization techniques.
Loading preview...
Model Overview
The cosmos1030/gmp-kd3e-1-s50pct-lr5e-5_20260828_125227 is a 4 billion parameter language model built upon the Qwen3 architecture. Its training incorporated a unique methodology, utilizing 50% unstructured data and a capped Proximal Gradient Descent (PGD) with a reverse-KL Optimal Policy Distribution (OPD) objective.
Key Characteristics
- Architecture: Qwen3-4B base model.
- Training Objective: Employs a specific capped PGD with reverse-KL OPD objective, indicating a focus on advanced optimization and policy distribution alignment during training.
- Data Mix: Trained with a significant portion (50%) of unstructured data, suggesting potential robustness to diverse input formats.
Potential Use Cases
This model is particularly interesting for researchers and developers exploring:
- Advanced Training Techniques: Investigating the impact of capped PGD and reverse-KL OPD objectives on model performance and behavior.
- Unstructured Data Processing: Applications requiring models trained on a high proportion of unstructured data.
- Experimental Deployments: As a base for further fine-tuning or research into specific language understanding or generation tasks, especially where the unique training methodology might offer advantages.