OpenRubrics/RubricRM-8B-Rubric
OpenRubrics/RubricRM-8B-Rubric is an 8 billion parameter RubricRM-Judge model, fine-tuned from Qwen3/Qwen3-8B. This model specializes in extracting universal, topic-agnostic rubric-style instructions from user requests, categorizing them into 'Hard Rule' and 'Principle' types. It is designed to generate comprehensive, concise, and unique evaluation criteria for assessing LLM responses. Its primary strength lies in creating structured rubrics for reward modeling and LLM alignment.
Loading preview...
OpenRubrics/RubricRM-8B-Rubric Overview
OpenRubrics/RubricRM-8B-Rubric is an 8 billion parameter language model specifically fine-tuned as a RubricRM-Judge. It is built upon the robust Qwen3/Qwen3-8B architecture and is designed to automate the generation of evaluation rubrics from natural language requests.
Key Capabilities
- Rubric Extraction: Automatically identifies and extracts rubric-style instructions from user prompts.
- Categorization: Distinguishes between two types of rubrics:
- Hard Rule: Strict requirements derived from explicit statements in the request (e.g., format, length).
- Principle: Abstract, domain-agnostic quality criteria (e.g., clarity, correctness, sound reasoning).
- Universality: Ensures all generated rubric items are universal principles, free from topic-specific references.
- Comprehensiveness: Aims to cover all critical aspects implied by the request, including explicit requirements and implicit quality standards.
- Conciseness & Uniqueness: Merges overlapping criteria and ensures precise, non-repetitive wording.
- Structured Output: Formats rubrics as a numbered list, with each item starting "The response" and appending
[Hard Rule]or[Principle].
Use Cases
This model is particularly well-suited for:
- Reward Modeling: Generating synthetic rubrics for training reward models in LLM alignment.
- Automated Evaluation: Creating objective evaluation criteria for assessing the quality of LLM-generated content.
- Instruction Following: Helping to define clear, measurable standards for how well an LLM adheres to complex instructions.
For more technical details and the underlying research, consider citing the associated paper: OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment.