IAAR-Shanghai/xVerify-3B-Ia
IAAR-Shanghai/xVerify-3B-Ia is a 3.2 billion parameter Llama-based model developed by IAAR-Shanghai, specifically fine-tuned as an evaluation tool for objective questions with single correct answers. It excels at extracting final answers from complex reasoning processes and efficiently identifying equivalence across various expression forms, including mathematical and natural language. This model is designed for broad applicability in evaluating tasks like math problems, multiple-choice questions, and classification.
Loading preview...
xVerify-3B-Ia: An Efficient Answer Verifier
xVerify-3B-Ia is a 3.2 billion parameter Llama-based model developed by IAAR-Shanghai, specifically designed as an evaluation tool for objective questions that have a single correct answer. Introduced in the paper "xVerify: Efficient Answer Verifier for Reasoning Model Evaluations" (arXiv:2504.10481), this model focuses on accurately extracting final answers from potentially lengthy reasoning chains and determining equivalence across different forms of expressions.
Key Capabilities
- Broad Applicability: Suitable for evaluating various objective question types, including mathematical problems, multiple-choice questions, classification tasks, and short-answer questions.
- Handles Long Reasoning Chains: Efficiently processes and extracts final answers from responses that involve extensive reasoning steps.
- Multilingual Support: Primarily supports Chinese and English responses, with compatibility for other languages.
- Powerful Equivalence Judgment: Features advanced capabilities for recognizing equivalence, such as:
- Basic transformations (e.g., letter case, Greek letter conversions).
- Mathematical equivalence across formats (LaTeX, fractions, scientific notation).
- Semantic equivalence in natural language answers.
- Matching multiple-choice responses by content rather than just option identifiers.
When to Use This Model
xVerify-3B-Ia is ideal for developers and researchers who need to automate the evaluation of reasoning models, particularly for tasks requiring precise answer verification against objective criteria. Its ability to handle complex reasoning and diverse answer formats makes it a robust tool for ensuring evaluation accuracy in educational, scientific, and technical domains.