FRPO/qwen3-1.7b-a15_global_token_norm-k1-cNone-globalTokNorm-clip0.2-mb4-eta100-bs256x5-n2-seed2

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026Architecture:Transformer Featherless Exclusive Cold

FRPO/qwen3-1.7b-a15_global_token_norm-k1-cNone-globalTokNorm-clip0.2-mb4-eta100-bs256x5-n2-seed2 is a 2 billion parameter language model based on the Qwen3-1.7B architecture, fine-tuned using Reinforcement Learning (RL) with the FRPO algorithm. It features a 32768 token context length and is specifically a checkpoint from KL-in-LLM-RL experiments. This model is designed for applications benefiting from RL-tuned performance, particularly within the context of the FRPO experimental framework.

Loading preview...

Model Overview

This model, FRPO/qwen3-1.7b-a15_global_token_norm-k1-cNone-globalTokNorm-clip0.2-mb4-eta100-bs256x5-n2-seed2, is a 2 billion parameter language model derived from the Qwen/Qwen3-1.7B base architecture. It has been fine-tuned using Reinforcement Learning (RL) with the FRPO algorithm as part of the KL-in-LLM-RL experimental series, utilizing the verl framework.

Key Characteristics

  • Base Model: Qwen3-1.7B, a causal language model.
  • Fine-tuning Method: Reinforcement Learning (RL) using the FRPO algorithm.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Weights: Provided in fp32 safetensors format, exactly as saved during training.
  • Experimental Context: Represents a specific checkpoint (global_step_200) from the KL-in-LLM-RL experiments, with its run configuration encoded in the repository name.

Intended Use Cases

This model is particularly suited for:

  • Research and Development: Exploring the effects of FRPO-based Reinforcement Learning on Qwen3-1.7B.
  • Comparative Analysis: Benchmarking RL-tuned models against their base counterparts or other RL algorithms.
  • Applications requiring RL-enhanced performance: Where the specific tuning methodology of FRPO is beneficial.