FRPO/qwen3-1.7b-a14_shuffle-k1-cNone-shuf-clip0.2-mb4-eta100-bs256x5-n2-seed2
FRPO/qwen3-1.7b-a14_shuffle-k1-cNone-shuf-clip0.2-mb4-eta100-bs256x5-n2-seed2 is a 2 billion parameter language model based on the Qwen3-1.7B architecture, fine-tuned using Reinforcement Learning (RL) with the FRPO (Featherless Reinforcement Learning Policy Optimization) method. This checkpoint is part of the KL-in-LLM-RL experiments, utilizing the verl framework. It is designed for tasks benefiting from RL fine-tuning, offering a specialized approach to language generation and understanding.
Loading preview...
Model Overview
FRPO/qwen3-1.7b-a14_shuffle-k1-cNone-shuf-clip0.2-mb4-eta100-bs256x5-n2-seed2 is a 2 billion parameter language model derived from the Qwen/Qwen3-1.7B base model. This specific checkpoint, global_step_200, represents a result from the KL-in-LLM-RL / FRPO experimental series, where it underwent Reinforcement Learning (RL) fine-tuning using the verl framework.
Key Characteristics
- Base Architecture: Qwen3-1.7B, a 2 billion parameter model.
- Fine-tuning Method: Utilizes the FRPO (Featherless Reinforcement Learning Policy Optimization) method within the KL-in-LLM-RL experiments.
- Weights: Provided in fp32 safetensors format, directly as saved by the trainer without post-processing.
- Configuration: The specific run configuration is embedded within the repository name, detailing parameters like shuffle, clip, and batch size.
Intended Use Cases
This model is particularly suited for research and development in:
- Exploring the effects of Reinforcement Learning fine-tuning on large language models.
- Applications where specialized RL-driven optimization of language generation is beneficial.
- Experiments requiring a Qwen3-1.7B variant with specific FRPO-based training characteristics.