rubricreward/mR3-Qwen3-4B-tgt-prompt-tgt-thinking

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 19, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

rubricreward/mR3-Qwen3-4B-tgt-prompt-tgt-thinking is a 4 billion parameter reward model, fine-tuned from Qwen3-4B. It is part of the mR3 family, designed for multilingual rubric-agnostic reward reasoning across 72 languages. This model excels at evaluating responses based on given rubrics, providing scores and reasoning for tasks like classification, preference optimization, and question answering.

Loading preview...

Model Overview

mR3-Qwen3-4B-tgt-prompt-tgt-thinking is a 4 billion parameter reward model, fine-tuned from the Qwen/Qwen3-4B base model. It is a key component of the mR3 (Multilingual Rubric-Agnostic Reward Reasoning Models) family, developed by rubricreward. This model is specifically designed to evaluate and provide reasoning for responses based on predefined rubrics.

Key Capabilities

  • Multilingual Reward Reasoning: Trained on a curated mR3 dataset covering 72 languages, enabling it to perform reward reasoning across a broad linguistic spectrum.
  • Rubric-Agnostic Evaluation: Capable of assessing responses for various tasks, including classification, preference optimization, and question answering, by considering an instruction, task description, input, response(s), evaluation rubrics, and a score with corresponding reasoning.
  • Detailed Evaluation Output: Each evaluation provides a score and a detailed explanation of the reasoning, available in both English and non-English languages.
  • Thinking Mode: Supports an 'enable_thinking' mode during generation, allowing for more structured and reasoned outputs.

When to Use This Model

  • Automated Content Evaluation: Ideal for scenarios requiring automated assessment of generated text based on specific criteria or rubrics.
  • Preference Optimization: Useful in reinforcement learning from human feedback (RLHF) pipelines for ranking and selecting preferred model responses.
  • Multilingual Applications: Particularly strong for applications needing evaluation capabilities across a wide range of languages, as it was trained on a 72-language dataset.
  • Research in Reward Modeling: A valuable tool for researchers exploring multilingual reward modeling and rubric-based evaluation systems. More details can be found in their paper.