FRPO/qwen3-1.7b-a8_klreward-krew-k1-coef1e-3-mb4-eta100-bs256x5-n2

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026Architecture:Transformer Featherless Exclusive Cold

FRPO/qwen3-1.7b-a8_klreward-krew-k1-coef1e-3-mb4-eta100-bs256x5-n2 is a 1.7 billion parameter language model based on the Qwen3 architecture, developed by FRPO. This model is an RL fine-tuned checkpoint from KL-in-LLM-RL / FRPO experiments, specifically optimized using the verl framework. It features a 32768 token context length and is designed for tasks benefiting from reinforcement learning fine-tuning.

Loading preview...

Model Overview

FRPO/qwen3-1.7b-a8_klreward-krew-k1-coef1e-3-mb4-eta100-bs256x5-n2 is a 1.7 billion parameter language model derived from the Qwen3-1.7B base model. It represents an RL fine-tuned checkpoint, specifically developed as part of the KL-in-LLM-RL / FRPO experimental series. The fine-tuning process utilized the verl framework, indicating an optimization approach focused on reinforcement learning.

Key Characteristics

  • Base Architecture: Qwen3-1.7B, providing a robust foundation for language understanding and generation.
  • Parameter Count: 1.7 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens, enabling processing of longer inputs and maintaining coherence over extended conversations or documents.
  • Fine-tuning Method: RL fine-tuned using the KL-in-LLM-RL / FRPO experimental setup, suggesting specialized capabilities in areas where reinforcement learning provides an advantage.
  • Weights: The model weights are provided in fp32 safetensors format, directly as saved by the trainer without additional post-processing.

Potential Use Cases

This model is suitable for applications requiring a compact yet capable language model that has undergone specific reinforcement learning optimization. Its large context window makes it valuable for tasks involving extensive text analysis, summarization, or generation where maintaining long-range dependencies is crucial. The RL fine-tuning implies potential strengths in areas like dialogue systems, interactive agents, or tasks benefiting from iterative refinement and reward-based learning.