CodeGoat24/UnifiedReward-Think-qwen3vl-4b

VISIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Nov 22, 2025License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

CodeGoat24/UnifiedReward-Think-qwen3vl-4b is a 4 billion parameter multimodal CoT reward model developed by CodeGoat24, based on the Qwen3VL architecture. It is uniquely designed for multi-dimensional, step-by-step long-chain reasoning across both visual understanding and generation reward tasks. This model specializes in evaluating complex multimodal reasoning processes, offering a unified approach to reward modeling. Its primary strength lies in its ability to assess intricate, sequential reasoning in multimodal contexts.

Loading preview...

Model Overview

UnifiedReward-Think-qwen3vl-4b is a 4 billion parameter model developed by CodeGoat24, representing the first unified multimodal Chain-of-Thought (CoT) reward model. It is specifically engineered to perform multi-dimensional, step-by-step long-chain reasoning, making it adept at evaluating complex processes in both visual understanding and generation reward tasks.

Key Capabilities

  • Unified Multimodal CoT Reward: Integrates reasoning across visual and generative modalities within a single reward framework.
  • Long-Chain Reasoning: Capable of evaluating and understanding sequential, multi-step thought processes.
  • Visual Understanding Reward: Assesses the quality of reasoning in tasks involving visual data interpretation.
  • Generation Reward: Provides feedback on the reasoning quality behind generated content.

When to Use This Model

This model is particularly well-suited for applications requiring sophisticated evaluation of multimodal AI systems, especially where the reasoning process itself needs to be assessed. It can be used to:

  • Develop more robust multimodal agents by providing nuanced reward signals.
  • Evaluate the step-by-step reasoning capabilities of other large language models in multimodal contexts.
  • Research into advanced reward modeling for complex AI tasks involving both vision and language.

For more technical details, refer to the accompanying paper and the project page.