ryokamoi/Qwen-2.5-7B-FoVer-PRM-old
The ryokamoi/Qwen-2.5-7B-FoVer-PRM-old is a 7.6 billion parameter Process Reward Model (PRM) developed by Ryo Kamoi and collaborators, based on the Qwen 2.5 architecture with a 32K context length. This model is specifically trained using formal verification tools (Z3, Isabelle) to provide step-level feedback on reasoning generated by large language models. It excels at identifying errors in formal logic and proof tasks, demonstrating cross-task transfer of verification capabilities to improve reasoning across mathematics, academic problems, and abstract logic.
Loading preview...
FoVer: Formal Verification for Process Reward Models
This model, ryokamoi/Qwen-2.5-7B-FoVer-PRM-old, is a 7.6 billion parameter Process Reward Model (PRM) based on the Qwen 2.5 architecture, developed by Ryo Kamoi and collaborators. It is designed to provide step-by-step feedback on the reasoning processes of large language models (LLMs).
Key Capabilities
- Automated Error Annotation: Utilizes formal verification tools (like Z3 and Isabelle) to automatically annotate step-level errors in LLM reasoning.
- Cross-Task Verification Transfer: PRMs trained on the FoVer dataset demonstrate improved verification capabilities across diverse reasoning tasks, including mathematics, academic problems, logic, and abstract reasoning.
- Step-Level Feedback: Provides granular feedback on individual steps within an LLM's solution, crucial for reinforcement learning and inference-time refinement.
- Specialized Dataset: Trained on the FoVer dataset, which includes automatically annotated step-level error labels for formal logic and proof tasks.
Good for
- Evaluating LLM Reasoning: Assessing the correctness of step-by-step reasoning in LLMs, particularly in formal logic and proof domains.
- Enhancing LLM Performance: Serving as a reward model for fine-tuning LLMs via reinforcement learning to improve their reasoning accuracy.
- Research in Formal Verification: Exploring the application of formal verification techniques to improve AI reasoning and trustworthiness.
This model represents a previous version of the FoVer project; users are encouraged to refer to the latest version for the most up-to-date materials and models.