CelineHuangxy/ICPO-Qwen3-8B-math-RS

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ICPO-Qwen3-8B-math-RS by CelineHuangxy is an 8 billion parameter language model based on the Qwen3 architecture, fine-tuned using the In-Context Steered Policy Optimization (ICPO) method with reward shaping. This model is specifically designed to enhance reasoning capabilities, particularly in mathematical tasks, by improving exploration and stability in Reinforcement Learning from Verifiable Rewards (RLVR). It leverages in-context learning to provide expert guidance without relying on external advanced models, making it suitable for complex mathematical reasoning problems.

Loading preview...

Overview

ICPO-Qwen3-8B-math-RS is an 8 billion parameter model developed by CelineHuangxy, based on the Qwen3 architecture. It implements the In-Context Steered Policy Optimization (ICPO) approach, which is a Reinforcement Learning with Verifiable Rewards (RLVR) method. This specific variant, ICPO-Qwen3-8B-math-RS, incorporates reward shaping to further stabilize optimization and enhance performance.

Key Capabilities

  • Enhanced Reasoning: Designed to improve the reasoning capabilities of Large Reasoning Models (LRMs), particularly in mathematical domains.
  • Improved Exploration: Addresses the limited exploration of existing RLVR methods by expanding trajectory diversity beyond the current policy's distribution.
  • In-Context Steered Policy Optimization: Utilizes the inherent in-context learning of LRMs to provide expert guidance, reducing reliance on computationally expensive advanced models.
  • Stabilized Training: Integrates expert region reject sampling to filter unreliable off-policy trajectories and annealed expert-bonus reward shaping to balance early guidance with autonomous improvement.

Use Cases

This model is particularly well-suited for:

  • Mathematical Reasoning Benchmarks: Demonstrates consistent enhancement in RLVR performance and training stability on mathematical reasoning tasks.
  • Research in RLVR: Offers a scalable and effective paradigm for researchers working on Reinforcement Learning from Verifiable Rewards for LRMs.
  • Complex Problem Solving: Applicable to scenarios requiring robust and verifiable reasoning, especially where exploration and stability are critical.