FRPO/qwen3-1.7b-a1_base-k1-cNone-clip0.2-mb4-eta100-bs256x5-n2

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026Architecture:Transformer Featherless Exclusive Cold

FRPO/qwen3-1.7b-a1_base-k1-cNone-clip0.2-mb4-eta100-bs256x5-n2 is a 2 billion parameter language model based on the Qwen3-1.7B architecture. Developed through KL-in-LLM-RL / FRPO experiments using the verl framework, this model is an RL fine-tuned checkpoint. It is specifically designed for research and development in reinforcement learning applications for large language models, offering a base for further experimentation.

Loading preview...

Model Overview

This model, FRPO/qwen3-1.7b-a1_base-k1-cNone-clip0.2-mb4-eta100-bs256x5-n2, is a 2 billion parameter language model derived from the Qwen/Qwen3-1.7B base architecture. It represents an RL fine-tuned checkpoint from the KL-in-LLM-RL / FRPO experimental series, utilizing the verl framework for its training.

Key Characteristics

  • Base Model: Qwen3-1.7B, providing a robust foundation.
  • Training Method: Reinforcement Learning (RL) fine-tuned, specifically within the KL-in-LLM-RL / FRPO experimental setup.
  • Weights: Provided in fp32 safetensors format, directly as saved by the trainer without additional post-processing.
  • Configuration: The run configuration details are embedded within the model's repository name.
  • Context Length: Supports a context length of 32768 tokens.

Intended Use Cases

This model is primarily suited for:

  • RL Research: Experimentation and development in reinforcement learning applied to large language models.
  • Comparative Studies: Serving as a baseline or comparison point for other RL-tuned models.
  • Fine-tuning: As a starting point for further specialized fine-tuning tasks where an RL-tuned base is beneficial.