double7/Qwen2.5-7B-GRRM

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Dec 29, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

double7/Qwen2.5-7B-GRRM is a 7.6 billion parameter Generative Reward Model (GenRM) developed by NJUNLP, specifically optimized for Machine Translation (MT). It implements the Group Quality Metric (GQM) paradigm, evaluating groups of candidate translations to produce relative rankings and reward scores. Initialized from Qwen2.5-7B, it was fine-tuned on Chinese-English data and optimized using Reinforcement Learning with Verifiable Rewards (RLVR) for high ranking accuracy. This model excels at providing relative rewards for MT outputs, particularly for Group-based Policy Optimization (GRPO) and pairwise/groupwise MT evaluation.

Loading preview...

Qwen2.5-7B-GRRM: A Group Relative Reward Model for Machine Translation

Qwen2.5-7B-GRRM is a 7.6 billion parameter Generative Reward Model (GenRM) developed by NJUNLP, designed to evaluate and rank machine translation (MT) outputs. It is built upon the Qwen2.5-7B architecture and specializes in the Group Quality Metric (GQM) paradigm, assessing a group of candidate translations simultaneously to generate relative rankings and reward scores.

Key Capabilities & Features

  • Group-based Evaluation: Evaluates up to 4 candidate translations within a group, providing relative rewards rather than absolute scores.
  • Optimized for MT: Specifically trained for machine translation quality assessment, particularly for use in Group-based Policy Optimization (GRPO).
  • Multilingual Support: While primarily trained on Chinese-English data, it demonstrates generalization across multiple languages including Portuguese, Spanish, French, German, Dutch, Italian, Korean, and Russian.
  • Structured Output: Generates a comparative analysis, a clear ranking (e.g., A > C > B), and consistent integer scores (0-10) for each candidate.
  • Two-Stage Training: Utilizes Supervised Fine-Tuning (SFT) for initial instruction following, followed by Reinforcement Learning with Verifiable Rewards (RLVR) to optimize ranking accuracy.

Ideal Use Cases

  • Reward Model for GRPO: Directly applicable as a reward model within GRPO frameworks for MT policy optimization.
  • Pairwise/Groupwise MT Evaluation: Suitable for scenarios requiring comparative assessment of MT outputs, where relative quality is more important than absolute scores.
  • Research & Development: Useful for researchers exploring advanced MT evaluation metrics and reward modeling techniques.

Limitations

  • Candidate Count: Best performance is observed with 2-4 candidates per group; ranking sensitivity may degrade with more.
  • Relative Scoring: Rewards are calibrated for within-group comparison and should not be interpreted as globally comparable quality scores across different inputs.
  • Generalization: While multilingual, performance may vary on low-resource languages or specialized domains.