FRPO/qwen3-1.7b-a1_base-k1-cNone-clip0.2-mb4-eta100-bs256x5-n2-seed1
FRPO/qwen3-1.7b-a1_base-k1-cNone-clip0.2-mb4-eta100-bs256x5-n2-seed1 is a 1.7 billion parameter language model based on the Qwen3 architecture, specifically fine-tuned using Reinforcement Learning (RL) through the KL-in-LLM-RL / FRPO experiments. This model is derived from the Qwen/Qwen3-1.7B base model and features fp32 safetensors weights. Its primary differentiator lies in its RL fine-tuning approach, making it suitable for research and applications exploring advanced language model optimization techniques.
Loading preview...
Model Overview
FRPO/qwen3-1.7b-a1_base-k1-cNone-clip0.2-mb4-eta100-bs256x5-n2-seed1 is a 1.7 billion parameter language model built upon the Qwen3-1.7B base architecture. This checkpoint is a result of the KL-in-LLM-RL / FRPO experimental series, utilizing the verl framework for Reinforcement Learning (RL) fine-tuning.
Key Characteristics
- Base Model: Qwen/Qwen3-1.7B
- Parameter Count: Approximately 1.7 billion parameters
- Fine-tuning Method: Reinforcement Learning (RL) via KL-in-LLM-RL / FRPO experiments
- Weights: Stored in fp32 safetensors format, directly from the trainer without post-processing.
- Configuration: The specific run configuration is encoded within the repository name itself.
Intended Use Cases
This model is particularly relevant for:
- RL Research: Experimenting with and evaluating the effects of KL-in-LLM-RL / FRPO fine-tuning techniques.
- Comparative Analysis: Benchmarking RL-tuned models against their base counterparts or other fine-tuning methodologies.
- Advanced NLP Development: Exploring applications where RL-based optimization might offer performance advantages or unique behavioral characteristics.