cosmos1030/gmp-kd3e-1-s50pct-lr5e-5_20260828_125050
The cosmos1030/gmp-kd3e-1-s50pct-lr5e-5_20260828_125050 model is a 4 billion parameter language model based on the Qwen3-4B architecture. It was trained using 50% unstructured, uncapped PGD with a reverse-KL OPD objective, and features a 32768 token context length. This model is a result of specific experimental training configurations, making it distinct for research into advanced policy gradient methods.
Loading preview...
Model Overview
The cosmos1030/gmp-kd3e-1-s50pct-lr5e-5_20260828_125050 is a 4 billion parameter language model derived from the Qwen3-4B architecture. It was developed by cosmos1030 and trained with a specific experimental setup involving 50% unstructured, uncapped Policy Gradient Descent (PGD) combined with a reverse-KL On-Policy Distribution (OPD) objective. The training process completed 2048 steps, though evaluation encountered an issue after the first benchmark.
Key Characteristics
- Architecture: Based on Qwen3-4B.
- Parameter Count: 4 billion parameters.
- Context Length: Supports a context window of 32768 tokens.
- Training Methodology: Utilizes a unique training regimen with 50% unstructured, uncapped PGD and a reverse-KL OPD objective.
- Experimental Focus: This model represents an archived experimental run, primarily serving as a record of a specific training configuration rather than a fully evaluated general-purpose model.
Intended Use Cases
This model is primarily suited for:
- Research and Development: Ideal for researchers studying advanced policy gradient methods, particularly those involving unstructured PGD and reverse-KL OPD objectives.
- Comparative Analysis: Can be used to compare the effects of its specific training methodology against other models or training paradigms.
- Understanding Training Dynamics: Provides a snapshot of a completed experimental training run, useful for analyzing training logs and outcomes (though its evaluation data is limited).