FRPO/qwen3-1.7b-a8_klreward-krew-k1-coef1e-2-mb4-eta100-bs256x5-n2

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 15, 2026Architecture:Transformer Featherless Exclusive Cold

FRPO/qwen3-1.7b-a8_klreward-krew-k1-coef1e-2-mb4-eta100-bs256x5-n2 is a 2 billion parameter language model based on the Qwen3-1.7B architecture, fine-tuned using Reinforcement Learning (RL) with the KL-in-LLM-RL / FRPO experimental framework. Developed as part of RL fine-tuning experiments, this model focuses on leveraging specific RL techniques. It is designed for research and development in advanced language model fine-tuning methodologies, particularly those involving KL-reward mechanisms.

Loading preview...

Model Overview

This model, FRPO/qwen3-1.7b-a8_klreward-krew-k1-coef1e-2-mb4-eta100-bs256x5-n2, is a 2 billion parameter language model derived from the Qwen/Qwen3-1.7B base architecture. It represents a checkpoint from the KL-in-LLM-RL / FRPO experimental series, specifically fine-tuned using Reinforcement Learning (RL) techniques.

Key Characteristics

  • Base Model: Utilizes Qwen/Qwen3-1.7B as its foundation.
  • Fine-tuning Method: Fine-tuned with Reinforcement Learning, specifically within the KL-in-LLM-RL / FRPO experimental framework.
  • Training Framework: Developed using the verl library.
  • Checkpoint: The repository contains the global_step_200 checkpoint.
  • Weights: Provided in fp32 safetensors format, directly as saved by the trainer.
  • Configuration: The specific run configuration is embedded within the model's repository name.

Intended Use Cases

This model is primarily suited for:

  • RL Research: Investigating the effects and performance of KL-in-LLM-RL and FRPO fine-tuning methods.
  • Experimental Development: Exploring advanced RL techniques for language model optimization.
  • Comparative Analysis: Benchmarking against other RL fine-tuned models or base models to understand the impact of specific training parameters.