RUC-AIBOX/AEWM

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 23, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

RUC-AIBOX/AEWM is a 35.1 billion parameter world model built on Qwen3.5-35B-A3B, designed to judge and edit an agent's decisions rather than predicting tool responses. It integrates Action Judge (AJ) for classifying decision quality and State Revision (SR) for correcting noisy actions. This model significantly enhances long-horizon reasoning and tool use across Search, Terminal, and Software Engineering domains by preventing task-state contamination.

Loading preview...

AEWM: Agent-Editing World Model

AEWM (Agent-Editing World Model) is a 35.1 billion parameter model, based on Qwen3.5-35B-A3B, that acts as a world model to improve the reliability of long-horizon agents. Unlike traditional world models that predict tool responses, AEWM focuses on judging and editing an agent's proposed decisions to mitigate "task-state contamination"—where agents make unsupported assumptions or persist with outdated plans.

Key Capabilities

  • Action Judge (AJ): Classifies an agent's proposed decision as Critical, Exploratory, or Noisy, determining whether to retain or intervene.
  • State Revision (SR): For noisy decisions, SR generates revised reasoning and actions, directly editing the agent's continuation before execution.
  • EditAct Framework: Integrates AJ and SR into an agent's interaction loop, allowing for real-time judgment, editing, and execution of actions.
  • Domain Expertise: Optimized for tool use in Search, Terminal, and Software Engineering (SWE) environments.

Performance Highlights

AEWM demonstrates strong performance in judging actions and improving end-to-end agent success:

  • Achieves 70.5% overall macro-F1 on the AEWM Action Judge Benchmark, outperforming DeepSeek-V4-Pro by 10.6 percentage points.
  • When integrated via EditAct, it provides absolute gains of 3.2 to 6.7 percentage points in mean scores across six benchmarks (BrowseComp, DeepSearchQA, Terminal-Bench 2.0, SWE-bench Pro, Doc2Repo, NL2Repo) compared to the strongest baselines, using Qwen3.5 backbones.

When to Use

Use AEWM if you are developing agents that require robust, long-horizon reasoning and tool use, particularly in Search, Terminal, or Software Engineering tasks. It is ideal for scenarios where preventing an agent from executing flawed or irrelevant actions is critical for task success and efficiency. This model serves as the world model, working alongside a separate agent that proposes initial reasoning and actions.