FRPO/qwen3-1.7b-a15_global_token_norm-k1-cNone-globalTokNorm-clip0.2-mb4-eta100-bs256x5-n2

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026Architecture:Transformer Featherless Exclusive Cold

FRPO/qwen3-1.7b-a15_global_token_norm-k1-cNone-globalTokNorm-clip0.2-mb4-eta100-bs256x5-n2 is a 2 billion parameter language model based on the Qwen3-1.7B architecture, developed by FRPO. This model is an RL fine-tuned checkpoint from the KL-in-LLM-RL / FRPO experiments, specifically optimized using the verl framework. It is designed for applications requiring a model enhanced through reinforcement learning, offering a 32768 token context length.

Loading preview...

Model Overview

FRPO/qwen3-1.7b-a15_global_token_norm-k1-cNone-globalTokNorm-clip0.2-mb4-eta100-bs256x5-n2 is a 2 billion parameter language model derived from the Qwen/Qwen3-1.7B base model. It represents an RL fine-tuned checkpoint resulting from the KL-in-LLM-RL / FRPO experimental series, utilizing the verl framework for training.

Key Characteristics

  • Base Architecture: Qwen3-1.7B
  • Parameter Count: Approximately 2 billion parameters.
  • Context Length: Supports a context window of 32768 tokens.
  • Fine-tuning Method: Reinforcement Learning (RL) fine-tuned using the FRPO method within the KL-in-LLM-RL experiments.
  • Weights: Provided in fp32 safetensors format, directly as saved by the trainer without additional post-processing.

Intended Use Cases

This model is particularly suited for research and development in:

  • Exploring the effects of Reinforcement Learning (RL) on language model performance.
  • Applications where a model fine-tuned with the FRPO (Fictitious Play for Reward Optimization) algorithm is beneficial.
  • Scenarios requiring a Qwen3-1.7B variant with specific RL-driven behavioral adjustments.