FRPO/qwen3-1.7b-a14_shuffle-k1-cNone-shuf-clip0.2-mb4-eta100-bs256x5-n2-seed1

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026Architecture:Transformer Featherless Exclusive Cold

FRPO/qwen3-1.7b-a14_shuffle-k1-cNone-shuf-clip0.2-mb4-eta100-bs256x5-n2-seed1 is a 1.7 billion parameter language model based on the Qwen3 architecture, specifically an RL fine-tuned checkpoint from the KL-in-LLM-RL / FRPO experiments. This model was trained using the verl framework and is derived from the Qwen/Qwen3-1.7B base model. It features fp32 safetensor weights and has a context length of 32768 tokens, making it suitable for tasks benefiting from reinforcement learning optimization.

Loading preview...

Model Overview

FRPO/qwen3-1.7b-a14_shuffle-k1-cNone-shuf-clip0.2-mb4-eta100-bs256x5-n2-seed1 is a 1.7 billion parameter language model that has undergone Reinforcement Learning (RL) fine-tuning. It is built upon the Qwen/Qwen3-1.7B base model and is a product of the KL-in-LLM-RL / FRPO experimental series, utilizing the verl training framework.

Key Characteristics

  • Base Model: Qwen3-1.7B architecture.
  • Fine-tuning: RL fine-tuned checkpoint from specific KL-in-LLM-RL / FRPO experiments.
  • Training Framework: Developed using the verl framework.
  • Weights: Provided in fp32 safetensors format, directly as saved by the trainer without further processing.
  • Context Length: Supports a substantial context window of 32768 tokens.

Potential Use Cases

This model is particularly relevant for researchers and developers interested in:

  • Exploring the effects of Reinforcement Learning (RL) fine-tuning on Qwen3-1.7B.
  • Experimenting with models derived from the KL-in-LLM-RL / FRPO research.
  • Applications requiring a model with a 32K context length and RL-based optimizations.