FRPO/qwen3-1.7b-a9_dapo-dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n2-r2
FRPO/qwen3-1.7b-a9_dapo-dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n2-r2 is a 2 billion parameter language model based on the Qwen3-1.7B architecture. This model is a checkpoint from the KL-in-LLM-RL / FRPO experiments, fine-tuned using Reinforcement Learning (RL) with the verl framework. It is specifically designed for research into RL fine-tuning methods, offering a specific configuration encoded in its repository name. The model provides fp32 safetensors weights, exactly as saved by the trainer.
Loading preview...
Model Overview
FRPO/qwen3-1.7b-a9_dapo-dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n2-r2 is a 2 billion parameter language model derived from the Qwen/Qwen3-1.7B base model. It represents a specific checkpoint from the KL-in-LLM-RL / FRPO experimental series, which focuses on Reinforcement Learning (RL) fine-tuning techniques.
Key Characteristics
- Base Model: Qwen3-1.7B, a 2 billion parameter architecture.
- Fine-tuning Method: Utilizes Reinforcement Learning (RL) through the verl framework, specifically as part of the KL-in-LLM-RL / FRPO experiments.
- Weights: Provided in fp32 safetensors format, preserving the exact state from the trainer without post-processing.
- Configuration: The model's specific run configuration is embedded within its repository name, indicating its experimental nature.
Intended Use Cases
This model is primarily suited for:
- RL Research: Ideal for researchers and developers exploring Reinforcement Learning fine-tuning strategies for large language models.
- Experimental Analysis: Useful for analyzing the effects of specific RL configurations and parameters, as encoded in its name.
- Comparative Studies: Can serve as a baseline or comparison point within the KL-in-LLM-RL / FRPO experimental context.