opencompass/CompassVerifier-32B

TEXT GENERATIONPricing:Input $2.72 / Output $4.8Concurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 9, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

CompassVerifier-32B is a 32.8 billion parameter verifier model developed by OpenCompass, built upon the Qwen series architecture. It is designed for robust evaluation and outcome reward of LLM responses across diverse domains including math, knowledge, and general reasoning. This model excels at identifying correct, incorrect, or problematic answers, handling various answer formats, and demonstrating strong performance and robustness to different prompt styles, making it ideal for automated LLM evaluation and as a reward model in reinforcement learning.

Loading preview...

CompassVerifier-32B: A Robust LLM Verifier

CompassVerifier-32B is a 32.8 billion parameter model developed by OpenCompass, specifically designed as an accurate and robust verifier for evaluating Large Language Model (LLM) outputs and serving as an outcome reward model. Built on the Qwen series architecture, it demonstrates strong multi-domain competency across math, knowledge, and diverse reasoning tasks.

Key Capabilities

  • Multi-domain Verification: Proficient in evaluating responses across general reasoning, knowledge, math, and science domains.
  • Diverse Answer Type Handling: Capable of processing various answer formats, including multi-subproblems, formulas, and sequence answers.
  • Robust Error Identification: Effectively identifies abnormal, invalid, or long-reasoning responses, and is robust to different prompt styles.
  • High Performance: Achieves an 87.7 F1 score on the VerifierBench benchmark, outperforming general LLMs and other verifier models in its class.
  • Reinforcement Learning (RL) Reward Model: Demonstrates superior performance when used as a reward model in RL training, significantly improving the reasoning capabilities of base models on complex math benchmarks like AIME and MATH500.

Use Cases

  • Automated LLM Evaluation: Ideal for automatically assessing the correctness and quality of LLM-generated answers.
  • Reinforcement Learning from Human Feedback (RLHF): Can be integrated as a reward model to guide the training of LLMs, particularly for improving reasoning and problem-solving abilities.
  • Quality Assurance: Useful for filtering or flagging problematic LLM outputs in applications requiring high accuracy.