CodeGoat24/WorldReward-qwen35-9b
WorldReward-qwen35-9b is a 9 billion parameter reward model developed by CodeGoat24, based on the Qwen3.5 architecture, designed for camera-conditioned world models. It specializes in evaluating video quality and alignment with actions and appearance in simulated environments, achieving 77.63% agreement with human labels on action evaluation on the WorldReward-Bench dataset. This model is optimized for assessing the realism and correctness of generated video sequences based on given captions and actions, with a context length of 32768 tokens.
Loading preview...
WorldReward-qwen35-9b: Reward Modeling for Camera-Conditioned World Models
WorldReward-qwen35-9b is a 9 billion parameter reward model developed by CodeGoat24, specifically designed for evaluating camera-conditioned world models. This model excels at assessing the quality and alignment of generated video content with specified actions and visual appearance, making it a crucial component for training and validating world models.
Key Capabilities
- Video Quality Assessment: Evaluates generated video sequences against human preferences for action, appearance, and motion.
- High Agreement with Human Labels: Achieves 77.63% agreement with human labels on action evaluation, 81.32% on appearance, and 73.03% on motion on the WorldReward-Bench dataset.
- Comparative Performance: Outperforms models like GPT-5.5, Gemini-3.1-Pro, and various Qwen3.5-VL variants in action and motion evaluation.
- Camera-Conditioned Evaluation: Specialized in scenarios where video generation is conditioned on camera parameters and actions.
Good For
- Training World Models: Providing reward signals for reinforcement learning in camera-conditioned world model development.
- Evaluating Video Generation: Benchmarking and assessing the realism and correctness of video outputs from generative models.
- Research in Embodied AI: Advancing research in agents that interact with and learn from simulated or real-world visual environments.