FRPO/qwen3-4b-a9_dapo-dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026Architecture:Transformer Featherless Exclusive Cold

FRPO/qwen3-4b-a9_dapo-dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4 is a 4 billion parameter language model based on the Qwen3 architecture, fine-tuned using Reinforcement Learning (RL) from the KL-in-LLM-RL / FRPO experiments. This model utilizes a Dapo-Dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4 configuration for its RL training. It is specifically designed for applications benefiting from RL-tuned responses, offering a 32768 token context length.

Loading preview...

Overview

This model, FRPO/qwen3-4b-a9_dapo-dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4, is a 4 billion parameter language model built upon the Qwen3-4B base architecture. It has undergone Reinforcement Learning (RL) fine-tuning as part of the KL-in-LLM-RL / FRPO experimental series, utilizing the verl framework.

Key Characteristics

  • Base Model: Qwen/Qwen3-4B
  • Parameter Count: 4 billion
  • Context Length: 32768 tokens
  • Fine-tuning Method: Reinforcement Learning (RL) with a specific Dapo-Dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4 configuration.
  • Weights: Provided in fp32 safetensors format, directly from the trainer without post-processing.

Intended Use Cases

This model is suitable for developers exploring the impact of specific RL fine-tuning strategies on Qwen3-4B. Its design suggests potential applications in scenarios where responses benefit from the particular Dapo-Dapo-noKL-clip0.2_0.28-mb4-eta100-bs256x5-n4 RL configuration, offering a substantial context window for complex tasks.