trillionlabs/gWorld-8B
trillionlabs/gWorld-8B is an 8 billion parameter GUI world model developed by Trillion Labs, based on the Qwen3-VL-8B architecture, designed for mobile GUI agents. It predicts the next screen state after a user action by generating renderable HTML/web code, rather than pixels, ensuring text legibility and structural accuracy. This model excels at action-conditioned next-state prediction for mobile GUIs, establishing a new Pareto frontier in GUI world modeling accuracy and outperforming larger models on GUI-specific benchmarks.
Loading preview...
Overview
trillionlabs/gWorld-8B is an 8 billion parameter vision-language model developed by Trillion Labs, accepted to ICML 2026. It functions as a GUI world model for mobile GUI agents, predicting the subsequent screen state after a user action. Unlike pixel-generation models, gWorld-8B generates renderable HTML/web code, which maintains text legibility and structural accuracy while avoiding common hallucination issues. The model is built upon the Qwen3-VL-8B architecture and processes a current screenshot and an action to output reasoning and renderable HTML.
Key Capabilities
- Generative Visual Code: Predicts next GUI states by generating executable HTML/CSS, ensuring sharp text and responsive layouts with a render failure rate of less than 1%.
- Action-Conditioned Prediction: Interprets user actions (e.g., TAP, TYPE) within a normalized coordinate space ([0, 1000]) and generates a "Next State Reasoning" block before the HTML output to ensure logical visual transitions.
- Efficiency and Accuracy: Establishes a new Pareto frontier, outperforming models up to 50.25x larger on GUI-specific benchmarks and achieving a +45.7% gain in Instruction Accuracy (IAcc.) over its base model.
- Zero-Shot Generalization: Demonstrates high performance on out-of-distribution benchmarks such such as AndroidWorld and KApps (Korean).
Good For
- Developing mobile GUI agents that require accurate next-state prediction.
- Applications needing to simulate mobile interface interactions with high fidelity.
- Research in GUI world modeling, especially for generative visual code approaches.