CelineHuangxy/ICPO-Qwen3-1.7B-math
CelineHuangxy/ICPO-Qwen3-1.7B-math is a 2 billion parameter Qwen3-based language model fine-tuned using the In-Context Steered Policy Optimization (ICPO) approach. This model is specifically designed to enhance reasoning capabilities, particularly in mathematical tasks, by leveraging in-context learning for expert guidance. It improves upon existing Reinforcement Learning from Verifiable Rewards (RLVR) methods by expanding exploration and stabilizing optimization without relying on advanced, inaccessible expert models. The model is optimized for robust performance on mathematical reasoning benchmarks.
Loading preview...
ICPO-Qwen3-1.7B-math: Enhanced Mathematical Reasoning
This model, ICPO-Qwen3-1.7B-math, is a 2 billion parameter variant of the Qwen3 architecture, fine-tuned using the novel In-Context Steered Policy Optimization (ICPO) framework. ICPO is an advanced Reinforcement Learning with Verifiable Rewards (RLVR) approach designed to significantly improve the reasoning capabilities of Large Reasoning Models (LRMs), especially in mathematical domains.
Key Differentiators & Capabilities
- In-Context Steered Policy Optimization (ICPO): Unlike traditional RLVR methods that rely on on-policy rollouts or expensive expert models, ICPO leverages the LRM's inherent in-context learning to provide expert guidance from existing datasets.
- Expanded Exploration: Introduces mixed-policy GRPO with implicit expert forcing, enabling broader exploration beyond the current policy distribution without needing advanced LRM trajectories.
- Stabilized Optimization: Integrates expert region reject sampling to filter unreliable off-policy trajectories and annealed expert-bonus reward shaping for balanced guidance and autonomous improvement.
- Mathematical Reasoning Focus: Consistently enhances RLVR performance and training stability on mathematical reasoning benchmarks, making it suitable for complex numerical and logical problems.
Use Cases
- Mathematical Problem Solving: Ideal for applications requiring accurate and robust solutions to mathematical questions.
- Reasoning Tasks: Suitable for general reasoning tasks where enhanced logical deduction is beneficial.
- Research in RLVR: Provides a scalable and effective paradigm for further research and development in Reinforcement Learning from Verifiable Rewards.