ryokamoi/Qwen-2.5-7B-FoVer-PRM-old

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 21, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The ryokamoi/Qwen-2.5-7B-FoVer-PRM-old is a 7.6 billion parameter Process Reward Model (PRM) developed by Ryo Kamoi and collaborators, based on the Qwen 2.5 architecture with a 32K context length. This model is specifically trained using formal verification tools (Z3, Isabelle) to provide step-level feedback on reasoning generated by large language models. It excels at identifying errors in formal logic and proof tasks, demonstrating cross-task transfer of verification capabilities to improve reasoning across mathematics, academic problems, and abstract logic.

Loading preview...

FoVer: Formal Verification for Process Reward Models

This model, ryokamoi/Qwen-2.5-7B-FoVer-PRM-old, is a 7.6 billion parameter Process Reward Model (PRM) based on the Qwen 2.5 architecture, developed by Ryo Kamoi and collaborators. It is designed to provide step-by-step feedback on the reasoning processes of large language models (LLMs).

Key Capabilities

  • Automated Error Annotation: Utilizes formal verification tools (like Z3 and Isabelle) to automatically annotate step-level errors in LLM reasoning.
  • Cross-Task Verification Transfer: PRMs trained on the FoVer dataset demonstrate improved verification capabilities across diverse reasoning tasks, including mathematics, academic problems, logic, and abstract reasoning.
  • Step-Level Feedback: Provides granular feedback on individual steps within an LLM's solution, crucial for reinforcement learning and inference-time refinement.
  • Specialized Dataset: Trained on the FoVer dataset, which includes automatically annotated step-level error labels for formal logic and proof tasks.

Good for

  • Evaluating LLM Reasoning: Assessing the correctness of step-by-step reasoning in LLMs, particularly in formal logic and proof domains.
  • Enhancing LLM Performance: Serving as a reward model for fine-tuning LLMs via reinforcement learning to improve their reasoning accuracy.
  • Research in Formal Verification: Exploring the application of formal verification techniques to improve AI reasoning and trustworthiness.

This model represents a previous version of the FoVer project; users are encouraged to refer to the latest version for the most up-to-date materials and models.