UKPLab/SciRM-Ref-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 14, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The UKPLab/SciRM-Ref-7B is a 7.6 billion parameter reward model developed by UKPLab, built upon Qwen2.5-7B-Instruct, specifically designed for evaluating scientific writing. It utilizes a two-stage reinforcement learning framework to optimize for scientific evaluation preferences and enhance reasoning capabilities through self-reflection. This model excels at multi-aspect evaluation, dynamic scoring rubrics, and cross-task generalization for diverse scientific writing tasks, offering a cost-efficient open-source solution.

Loading preview...

SciRM-Ref-7B: Reward Model for Scientific Writing Evaluation

SciRM-Ref-7B is a 7.6 billion parameter reward model developed by UKPLab, based on the Qwen2.5-7B-Instruct architecture. It is specifically engineered for the nuanced evaluation of scientific writing, employing a sophisticated two-stage reinforcement learning framework (using GRPO) to achieve its capabilities. The initial stage optimizes for general scientific evaluation preferences, while the second stage refines reasoning through self-reflection on previous answers.

Key Capabilities & Features

  • Multi-aspect evaluation: Assesses multiple dimensions of evaluation per task, providing fine-grained feedback.
  • Dynamic scoring rubrics: Adapts to explicit evaluation constitutions provided at both training and inference times, allowing for flexible criteria.
  • Cross-task generalization: Capable of handling diverse and previously unseen scientific writing tasks without requiring task-specific retraining.
  • Reasoning enhancement: The 'Ref' in SciRM-Ref signifies its enhanced reasoning capabilities compared to its predecessor, SciRM-7B, achieved via self-reflection.
  • Open-source and cost-efficient: Provides an accessible solution without reliance on proprietary large language models.

Use Cases & Considerations

This model is primarily designed for evaluating scientific writing, including tasks like related work section generation and scientific review writing. It demonstrates generalization to unseen tasks such as novelty evaluation alignment and paper revision evaluation. While SciRM-Ref-7B offers increased reasoning, it is crucial to provide a clear system prompt and explicit evaluation criteria for optimal performance. The model's training is primarily on English scientific text from NLP/CS domains, and performance may vary in other scientific fields or languages. It is recommended not to use this model as the sole arbiter for high-stakes scientific publishing decisions.