CodeGoat24/WorldReward-qwen35-9b

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 28, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

WorldReward-qwen35-9b is a 9 billion parameter reward model developed by CodeGoat24, based on the Qwen3.5 architecture, designed for camera-conditioned world models. It specializes in evaluating video quality and alignment with actions and appearance in simulated environments, achieving 77.63% agreement with human labels on action evaluation on the WorldReward-Bench dataset. This model is optimized for assessing the realism and correctness of generated video sequences based on given captions and actions, with a context length of 32768 tokens.

Loading preview...

WorldReward-qwen35-9b: Reward Modeling for Camera-Conditioned World Models

WorldReward-qwen35-9b is a 9 billion parameter reward model developed by CodeGoat24, specifically designed for evaluating camera-conditioned world models. This model excels at assessing the quality and alignment of generated video content with specified actions and visual appearance, making it a crucial component for training and validating world models.

Key Capabilities

  • Video Quality Assessment: Evaluates generated video sequences against human preferences for action, appearance, and motion.
  • High Agreement with Human Labels: Achieves 77.63% agreement with human labels on action evaluation, 81.32% on appearance, and 73.03% on motion on the WorldReward-Bench dataset.
  • Comparative Performance: Outperforms models like GPT-5.5, Gemini-3.1-Pro, and various Qwen3.5-VL variants in action and motion evaluation.
  • Camera-Conditioned Evaluation: Specialized in scenarios where video generation is conditioned on camera parameters and actions.

Good For

  • Training World Models: Providing reward signals for reinforcement learning in camera-conditioned world model development.
  • Evaluating Video Generation: Benchmarking and assessing the realism and correctness of video outputs from generative models.
  • Research in Embodied AI: Advancing research in agents that interact with and learn from simulated or real-world visual environments.