PKU-ONELab/CE-RM-4B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jan 31, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

PKU-ONELab/CE-RM-4B is a 4 billion parameter pointwise generative reward model developed by PKU-ONELab, optimized via a two-stage rollout method and unified query-based criteria. Trained on approximately 5.7K high-quality data from open-source preference datasets, this model excels in automatic evaluation for open-ended natural language generation. It achieves superior performance on diverse reward model benchmarks, particularly in Best-of-N scenarios, and demonstrates effective improvements in downstream reinforcement learning practices.

Loading preview...

Overview

PKU-ONELab/CE-RM-4B is a 4 billion parameter generative reward model designed for automatic evaluation of open-ended natural language generation. Developed by PKU-ONELab, this model addresses limitations in prior LLM-as-a-Judge paradigms, such as the dominance of pairwise evaluation and inadequate optimization of evaluation criteria.

Key Capabilities

  • Pointwise Generative Reward Modeling: Unlike traditional pairwise evaluation, CE-RM-4B focuses on pointwise assessment, providing more granular feedback.
  • Optimized Training: It is trained using a dedicated two-stage rollout method and unified query-based criteria, leveraging about 5.7K high-quality data from open-source preference datasets.
  • Superior Benchmark Performance: The model achieves strong results on various reward model benchmarks, particularly demonstrating effectiveness in Best-of-N scenarios.
  • Enhanced RL Practice: CE-RM-4B is shown to deliver more effective improvements in downstream reinforcement learning applications, bridging the gap between benchmark performance and practical utility.

When to Use This Model

  • Automatic Evaluation: Ideal for evaluating open-ended natural language generation where rule-based metrics are insufficient.
  • Reinforcement Learning from Human Feedback (RLHF): Suitable as a generative reward model to guide and improve the training of other language models.
  • Custom Criteria Generation: Capable of generating a minimal set of evaluation criteria based on a user query, and then using these criteria to evaluate responses.