ant-intl/O2-27b-preview

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

O2-27B-Preview is a 27 billion parameter agentic model developed by Ant International, initialized from Qwen3.5-27B, with a 32768 token context length. It is uniquely trained on interactions from real office and software-development workflows, learning from executed actions, observed consequences, and repairs. This model excels in agentic reasoning, tool use, and coding tasks, demonstrating significant improvements over its base model in preliminary evaluations.

Loading preview...

O2-27B-Preview: Agentic Model for Real-World Workflows

O2-27B-Preview is a 27 billion parameter model developed by Ant International, built upon Qwen3.5-27B, and designed for advanced agentic reasoning, tool use, and coding. Its core differentiator lies in its unique training methodology: it learns from interactions generated within live operational environments, including international merchant customer-service and software development workflows. This approach allows the model to learn not just from successful outcomes, but also from failures and subsequent repairs, capturing a more nuanced understanding of task execution.

Key Capabilities & Differentiators

  • Real-World Training: Trained by deploying an executable agent runtime in actual office and development environments, capturing interactions with tools, repositories, and workflow states.
  • Learning from Executable Outcomes: Utilizes task-conditioned verifier suites to check observable evidence (tool results, environment state, code execution), providing concrete feedback.
  • Preserving Failures and Repairs: Records task context, model actions, executed consequences, verifier outcomes, localized failure feedback, and repaired actions, enabling learning from the entire interaction trajectory.
  • Performance: Preliminary evaluations show O2-27B-Preview achieves an average score of 73.40 across eleven reasoning, tool-use, and coding benchmarks, improving upon its Qwen3.5-27B base model's 69.97 average.

Intended Use Cases

  • Agentic Reasoning: Multi-step tasks requiring complex decision-making.
  • Tool-Using Assistants: Operating in controlled environments with explicit validation.
  • Software Development: Workflows involving code generation, testing, and repository interactions.
  • Business Process Automation: Office and business automation with integrated validation.
  • Research: Exploring runtime learning, executable verification, and on-policy distillation.