CodeGoat24/WorldReward-qwen38-27b

Hugging Face
VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 6, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

CodeGoat24/WorldReward-qwen38-27b is a 27 billion parameter model developed by CodeGoat24, specifically designed for reward modeling in camera-conditioned world models. This model is based on the Qwen architecture and is optimized for evaluating and understanding video-based environmental interactions. It excels at processing visual inputs to provide reward signals, making it suitable for robotics and reinforcement learning applications.

Loading preview...

WorldReward-qwen38-27b Overview

CodeGoat24/WorldReward-qwen38-27b is a 27 billion parameter model focused on reward modeling for camera-conditioned world models. Developed by CodeGoat24, this model is part of the WorldReward project, which aims to provide robust evaluation mechanisms for agents interacting with visual environments.

Key Capabilities

  • Camera-Conditioned Reward Modeling: Specifically designed to process visual input from cameras and generate reward signals, crucial for training agents in complex environments.
  • Video Analysis: Capable of analyzing video sequences (e.g., my_data/system_x.mp4, my_data/system_y.mp4) in conjunction with a reference image and textual captions.
  • Action-Based Evaluation: Integrates action sequences (e.g., forward,forward,left+camera_down) to understand and evaluate the impact of actions within a visual context.
  • Reasoning Display: Supports displaying the reasoning behind its reward evaluations, enhancing transparency and interpretability.

Good For

  • Reinforcement Learning: Ideal for researchers and developers building and training reinforcement learning agents that operate in visually rich, dynamic environments.
  • Robotics: Applicable in robotics for tasks requiring an understanding of environmental feedback and goal-oriented behavior based on camera input.
  • World Model Development: Useful for those developing and evaluating world models that simulate and predict future states based on visual observations and actions.

For more technical details, refer to the WorldReward paper and the project page.