UKPLab/SciRM-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 14, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The UKPLab/SciRM-7B model, developed by UKPLab, is a 7 billion parameter reward model based on Qwen2.5-7B-Instruct, specifically designed for evaluating scientific writing. It utilizes a two-stage reinforcement learning framework to assess multiple evaluation dimensions and adapt to dynamic scoring rubrics. This model excels at cross-task generalization, handling diverse scientific writing tasks without task-specific retraining, making it a cost-efficient open-source solution for scientific text evaluation.

Loading preview...

Overview

UKPLab/SciRM-7B is an open-source, cost-efficient reward model developed by UKPLab for the evaluation of scientific writing. Built upon the Qwen2.5-7B-Instruct base model, SciRM-7B is trained using a two-stage reinforcement learning framework, specifically GRPO, to optimize scientific evaluation preferences. A related model, SciRM-Ref-7B, further enhances reasoning capabilities through self-reflection.

Key Capabilities

  • Multi-aspect evaluation: Assesses multiple dimensions of scientific writing quality per task.
  • Dynamic scoring rubrics: Conditions evaluation on explicit criteria at both training and inference times, allowing adaptability.
  • Cross-task generalization: Capable of handling diverse and previously unseen scientific writing tasks without requiring task-specific retraining.
  • Cost-efficient: Designed to be open-source and does not rely on proprietary large language models.
  • Trained on specific tasks: Evaluated on tasks like Related Work Section Generation and Scientific Review Writing, demonstrating generalization to unseen tasks such as Novelty Evaluation Alignment and Paper Revision Evaluation.

When to Use This Model

SciRM-7B is ideal for developers and researchers needing an automated system to evaluate scientific texts based on defined criteria. It is particularly useful for tasks involving detailed feedback generation, quality assessment of scientific articles, or educational tools for scientific writing. Users should provide a clear evaluation constitution and adhere to the system prompt for optimal performance. It is important to note that while it has increased reasoning capabilities, it is not a complete reasoning model and should not be the sole arbiter in high-stakes publishing decisions. The model is primarily trained on English scientific text from NLP/CS domains, and performance may vary in other fields or languages.