PKU-ONELab/CE-RM-4B
PKU-ONELab/CE-RM-4B is a 4 billion parameter pointwise generative reward model developed by PKU-ONELab, optimized via a two-stage rollout method and unified query-based criteria. Trained on approximately 5.7K high-quality data from open-source preference datasets, this model excels in automatic evaluation for open-ended natural language generation. It achieves superior performance on diverse reward model benchmarks, particularly in Best-of-N scenarios, and demonstrates effective improvements in downstream reinforcement learning practices.
Loading preview...
Overview
PKU-ONELab/CE-RM-4B is a 4 billion parameter generative reward model designed for automatic evaluation of open-ended natural language generation. Developed by PKU-ONELab, this model addresses limitations in prior LLM-as-a-Judge paradigms, such as the dominance of pairwise evaluation and inadequate optimization of evaluation criteria.
Key Capabilities
- Pointwise Generative Reward Modeling: Unlike traditional pairwise evaluation, CE-RM-4B focuses on pointwise assessment, providing more granular feedback.
- Optimized Training: It is trained using a dedicated two-stage rollout method and unified query-based criteria, leveraging about 5.7K high-quality data from open-source preference datasets.
- Superior Benchmark Performance: The model achieves strong results on various reward model benchmarks, particularly demonstrating effectiveness in Best-of-N scenarios.
- Enhanced RL Practice: CE-RM-4B is shown to deliver more effective improvements in downstream reinforcement learning applications, bridging the gap between benchmark performance and practical utility.
When to Use This Model
- Automatic Evaluation: Ideal for evaluating open-ended natural language generation where rule-based metrics are insufficient.
- Reinforcement Learning from Human Feedback (RLHF): Suitable as a generative reward model to guide and improve the training of other language models.
- Custom Criteria Generation: Capable of generating a minimal set of evaluation criteria based on a user query, and then using these criteria to evaluate responses.