XingYing-stack/TIPS-Qwen3-4B-Instruct-2507-Math
The XingYing-stack/TIPS-Qwen3-4B-Instruct-2507-Math is a 4 billion parameter generative reward model, initialized from Qwen/Qwen3-4B-Instruct-2507. It is specifically trained with outcome-only GRPO to reason over mathematical solutions, producing step-level and outcome labels. This model is designed for reward modeling and process verification in mathematical contexts, rather than general-purpose chat applications.
Loading preview...
Model Overview
The TIPS-Qwen3-4B-Instruct-2507-Math is a specialized 4 billion parameter generative reward model. It is built upon the Qwen/Qwen3-4B-Instruct-2507 base model and has undergone further training using an outcome-only Generative Reward Policy Optimization (GRPO) approach.
Key Capabilities
- Mathematical Reasoning: Designed to analyze and reason over mathematical solutions.
- Generative Reward Modeling: Produces both step-level and final outcome labels for mathematical problems.
- Process Verification: Intended for verifying the correctness of solution steps and overall outcomes in mathematical tasks.
Intended Use
This model is specifically purposed for:
- Reward modeling in mathematical domains.
- Verifying the process and results of mathematical solutions.
It is not recommended for general-purpose conversational AI or chat applications. Users should refer to the TIPS repository for appropriate prompt templates and evaluation scripts. Training data is available at XingYing-stack/TIPS-Training-Data.