FRPO/qwen3-1.7b-a11_lengthnorm_center-k1-cGroupBoth-lnorm-clip0.2-mb4-eta100-bs256x5-n2
FRPO/qwen3-1.7b-a11_lengthnorm_center-k1-cGroupBoth-lnorm-clip0.2-mb4-eta100-bs256x5-n2 is a 2 billion parameter language model based on the Qwen3-1.7B architecture, fine-tuned using Reinforcement Learning (RL) from the KL-in-LLM-RL / FRPO experiments. This model is specifically optimized through a unique RL training configuration, as indicated by its detailed naming convention. It is designed for applications benefiting from RL-enhanced language generation, offering a 32768 token context length.
Loading preview...
Overview
This model, FRPO/qwen3-1.7b-a11_lengthnorm_center-k1-cGroupBoth-lnorm-clip0.2-mb4-eta100-bs256x5-n2, is a 2 billion parameter language model derived from the Qwen/Qwen3-1.7B base architecture. It has undergone specific Reinforcement Learning (RL) fine-tuning as part of the KL-in-LLM-RL / FRPO experimental series, utilizing the verl framework.
Key Characteristics
- Base Model: Qwen3-1.7B, a compact yet capable foundation.
- Fine-tuning: Enhanced through a specialized RL process, indicated by the detailed configuration in its name (e.g.,
lengthnorm_center-k1-cGroupBoth-lnorm-clip0.2-mb4-eta100-bs256x5-n2). - Checkpoint: The primary checkpoint available is
global_step_200. - Weights: Provided in fp32 safetensors format, directly from the trainer without post-processing.
- Context Length: Supports a substantial context window of 32768 tokens.
Good For
- Researchers and developers exploring the impact of specific RL fine-tuning configurations on language models.
- Use cases requiring a Qwen3-1.7B-based model with performance characteristics shaped by advanced RL techniques.
- Applications where the detailed training parameters encoded in the model name are relevant for reproducibility or comparative analysis.