FRPO/qwen3-1.7b-a9_dapo-dapo-noKL-clip0.2_0.35-mb4-eta100-bs256x5-n2-r3

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026Architecture:Transformer Featherless Exclusive Cold

FRPO/qwen3-1.7b-a9_dapo-dapo-noKL-clip0.2_0.35-mb4-eta100-bs256x5-n2-r3 is a 2 billion parameter language model based on the Qwen3-1.7B architecture, fine-tuned using Reinforcement Learning (RL) with the FRPO experimental framework. This model, developed through KL-in-LLM-RL experiments, is optimized for specific RL-driven language generation tasks. It features a 32768 token context length and is suitable for applications requiring models trained with advanced RL techniques.

Loading preview...

Model Overview

This repository hosts FRPO/qwen3-1.7b-a9_dapo-dapo-noKL-clip0.2_0.35-mb4-eta100-bs256x5-n2-r3, a 2 billion parameter language model derived from the Qwen3-1.7B base architecture. It has been fine-tuned using Reinforcement Learning (RL) as part of the KL-in-LLM-RL / FRPO experimental series, utilizing the verl training framework.

Key Characteristics

  • Base Model: Qwen/Qwen3-1.7B, a robust foundation for language generation.
  • Fine-tuning Method: Leverages advanced Reinforcement Learning (RL) techniques within the FRPO experimental setup.
  • Parameter Count: Approximately 2 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens, enabling processing of longer inputs and generating more coherent, extended outputs.
  • Checkpoint: The primary checkpoint available is global_step_200, provided in fp32 safetensors format without post-processing.
  • Configuration: The specific run configuration used during training is encoded directly in the repository name.

Intended Use Cases

This model is particularly well-suited for research and development in:

  • RL-driven Language Generation: Ideal for experiments and applications requiring models fine-tuned with specific Reinforcement Learning objectives.
  • Advanced Fine-tuning Research: Useful for exploring the impact of different RL algorithms and configurations on language model performance.
  • Context-rich Applications: Its large context window makes it suitable for tasks demanding extensive contextual understanding and generation.