ant-intl/O2-9b-preview
O2-9B-Preview is a 9 billion parameter agentic model developed by Ant International, initialized from Qwen3.5-9B. It is uniquely trained from interactions generated within real office and software-development workflows, observing consequences and repairs. This model excels in agentic reasoning, tool use, and coding tasks, making it suitable for automating complex operational environments.
Loading preview...
O2-9B-Preview: An Agentic Model for Real-World Workflows
O2-9B-Preview is a 9-billion parameter model from Ant International, built upon Qwen3.5-9B. Its core innovation lies in its training methodology: it learns from interactions within live operational environments, including customer service and software development workflows. This allows the model to observe actions, consequences, and necessary repairs, moving beyond static prompt-response datasets.
Key Capabilities
- Agentic Reasoning: Designed to handle multi-step tasks by acting, observing, and repairing actions.
- Enhanced Tool Use: Trained to interact with tools, repositories, and workflow states in deployed settings.
- Coding Proficiency: Demonstrates improved performance in coding benchmarks compared to its base model.
- Learning from Failures: Records task context, model actions, executed consequences, verifier outcomes, and localized failure feedback, including repaired actions.
What Makes O2 Different
Unlike models trained solely on static data, O2-9B-Preview's training loop is connected to deployed environments. It learns from executable outcomes, where a task-conditioned verifier suite checks observable evidence. This process, called Runtime On-Policy Sequence Distillation (OPSD), distills student-visited states into the model, providing a grounded and renewable source of training experience.
Intended Use Cases
- Research and evaluation of agentic reasoning over multi-step tasks.
- Developing tool-using assistants for controlled environments.
- Automating software development and office processes with explicit validation.
- Research into runtime learning and executable verification.