ChengyuDu0123/HER-RM-32B

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 29, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

ChengyuDu0123/HER-RM-32B is a 32 billion parameter Generative Reward Model (GenRM) specifically designed for evaluating role-playing responses, developed by Chengyu Du et al. Unlike traditional reward models, HER-RM generates dynamic, by-case evaluation principles based on dialogue context and provides detailed comparative analysis. It excels at nuanced assessment of role-playing scenarios, identifying implicit human preferences across 12 evaluation dimensions. This model is optimized for providing in-depth, context-aware feedback for AI role-playing applications.

Loading preview...

HER-RM: Generative Reward Model for Role-Playing

HER-RM (Human-like Reasoning and Reinforcement Learning for LLM Role-playing Reward Model) is a 32 billion parameter Generative Reward Model (GenRM) developed by Chengyu Du et al. It is specifically engineered to evaluate the quality of responses in role-playing scenarios. Unlike conventional reward models that output a single scalar score, HER-RM generates by-case evaluation principles dynamically based on the dialogue context, offering detailed comparative analysis between two candidate responses.

Key Capabilities

  • Context-Aware Evaluation: Identifies implicit human preferences and generates relevant evaluation principles for each specific dialogue.
  • Detailed Analysis: Provides principle-by-principle comparisons with reasoning, determining a winner between two responses.
  • Comprehensive Dimensions: Evaluates responses across 12 dimensions, including 6 negative (e.g., Character OOC, Contradiction with context) and 6 positive (e.g., Character fidelity, Emotional authenticity).
  • Trained on Extensive Data: Leverages over 300,000 preference pairs from role-playing scenarios and 107 distilled evaluation principles.

Good For

  • Developing and refining AI role-playing agents: Provides granular feedback to improve model performance in persona simulation.
  • Automated evaluation of role-play dialogues: Offers a more nuanced and explainable assessment than traditional scalar reward models.
  • Research into human-like reasoning for LLM evaluation: Demonstrates a novel approach to cognitive-level persona simulation.