FRPO/qwen3-1.7b-a9_dapo-dapo-noKL-clip0.2-mb4-eta100-bs256x5-n2-r3

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026Architecture:Transformer Featherless Exclusive Cold

FRPO/qwen3-1.7b-a9_dapo-dapo-noKL-clip0.2-mb4-eta100-bs256x5-n2-r3 is a 2 billion parameter language model based on the Qwen3-1.7B architecture. Developed through FRPO experiments using the verl framework, this checkpoint is specifically fine-tuned with Reinforcement Learning (RL). It is designed for research into RL fine-tuning methods for large language models.

Loading preview...

Model Overview

This model, FRPO/qwen3-1.7b-a9_dapo-dapo-noKL-clip0.2-mb4-eta100-bs256x5-n2-r3, is a 2 billion parameter language model derived from the Qwen/Qwen3-1.7B base model. It represents a specific checkpoint from the KL-in-LLM-RL / FRPO experimental series, fine-tuned using Reinforcement Learning (RL) with the verl framework.

Key Characteristics

  • Base Model: Qwen3-1.7B, a 2 billion parameter architecture.
  • Fine-tuning Method: Reinforcement Learning (RL) via the FRPO (Featherless Reinforcement Learning Policy Optimization) experiments.
  • Framework: Utilizes the verl library for its RL training.
  • Weights: Provided in fp32 safetensors format, directly as saved from the trainer without post-processing.
  • Configuration: The specific run configuration is encoded within the repository name, detailing parameters like clip, mb, eta, bs, n, and r values.

Intended Use Cases

This model is primarily suited for:

  • RL Research: Investigating the effects and performance of specific RL fine-tuning configurations on language models.
  • Experimental Analysis: Studying the impact of parameters like clip0.2, mb4, eta100, bs256x5, n2, and r3 on model behavior and capabilities.
  • Comparative Studies: Benchmarking different RL fine-tuning approaches against this specific checkpoint.