ant-intl/O2-9b-preview

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

O2-9B-Preview is a 9 billion parameter agentic model developed by Ant International, initialized from Qwen3.5-9B. It is uniquely trained from interactions generated within real office and software-development workflows, observing consequences and repairs. This model excels in agentic reasoning, tool use, and coding tasks, making it suitable for automating complex operational environments.

Loading preview...

O2-9B-Preview: An Agentic Model for Real-World Workflows

O2-9B-Preview is a 9-billion parameter model from Ant International, built upon Qwen3.5-9B. Its core innovation lies in its training methodology: it learns from interactions within live operational environments, including customer service and software development workflows. This allows the model to observe actions, consequences, and necessary repairs, moving beyond static prompt-response datasets.

Key Capabilities

  • Agentic Reasoning: Designed to handle multi-step tasks by acting, observing, and repairing actions.
  • Enhanced Tool Use: Trained to interact with tools, repositories, and workflow states in deployed settings.
  • Coding Proficiency: Demonstrates improved performance in coding benchmarks compared to its base model.
  • Learning from Failures: Records task context, model actions, executed consequences, verifier outcomes, and localized failure feedback, including repaired actions.

What Makes O2 Different

Unlike models trained solely on static data, O2-9B-Preview's training loop is connected to deployed environments. It learns from executable outcomes, where a task-conditioned verifier suite checks observable evidence. This process, called Runtime On-Policy Sequence Distillation (OPSD), distills student-visited states into the model, providing a grounded and renewable source of training experience.

Intended Use Cases

  • Research and evaluation of agentic reasoning over multi-step tasks.
  • Developing tool-using assistants for controlled environments.
  • Automating software development and office processes with explicit validation.
  • Research into runtime learning and executable verification.