FRPO/qwen3-4b-a9_dapo-dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4
FRPO/qwen3-4b-a9_dapo-dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4 is a 4 billion parameter language model based on the Qwen3 architecture, fine-tuned using Reinforcement Learning (RL) from the KL-in-LLM-RL / FRPO experiments. This model utilizes a Dapo-Dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4 configuration for its RL training. It is specifically designed for applications benefiting from RL-tuned responses, offering a 32768 token context length.
Loading preview...
Overview
This model, FRPO/qwen3-4b-a9_dapo-dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4, is a 4 billion parameter language model built upon the Qwen3-4B base architecture. It has undergone Reinforcement Learning (RL) fine-tuning as part of the KL-in-LLM-RL / FRPO experimental series, utilizing the verl framework.
Key Characteristics
- Base Model: Qwen/Qwen3-4B
- Parameter Count: 4 billion
- Context Length: 32768 tokens
- Fine-tuning Method: Reinforcement Learning (RL) with a specific Dapo-Dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4 configuration.
- Weights: Provided in fp32 safetensors format, directly from the trainer without post-processing.
Intended Use Cases
This model is suitable for developers exploring the impact of specific RL fine-tuning strategies on Qwen3-4B. Its design suggests potential applications in scenarios where responses benefit from the particular Dapo-Dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4 RL configuration, offering a substantial context window for complex tasks.