BAAI/AREX-2

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:apache-2.0Architecture:Transformer0.1K Open Weights Featherless Exclusive Cold

BAAI/AREX-2 is a 27B-parameter long-horizon agent model developed by the Beijing Academy of Artificial Intelligence (BAAI), built on a Qwen3.8-compatible multimodal architecture. It specializes in self-improving solutions over multiple test-time rounds through feedback-driven reflection. The model excels in long-horizon reasoning for tasks like algorithmic programming, machine-learning engineering, and deep research, leveraging its 262,144 token context length.

Loading preview...

BAAI/AREX-2: Self-Improving Agent Model

BAAI/AREX-2 is a 27-billion parameter agent model developed by the Beijing Academy of Artificial Intelligence (BAAI). Built on a Qwen3.8-compatible multimodal architecture, it is designed for long-horizon tasks, featuring a substantial 262,144 token context length. The model's core innovation lies in its ability to self-improve solutions iteratively through a 'propose, measure, reflect, revise' cycle, driven by verifiable feedback such as scores, logs, errors, and timings.

Key Capabilities

  • Long-horizon self-improvement: Effectively refines solutions over multiple test-time rounds.
  • Feedback-driven reflection: Utilizes various forms of feedback to guide subsequent revisions.
  • Cross-domain performance: Training on coding and machine-learning tasks enhances its deep-research capabilities.
  • Sustained reasoning: Maintains productive iteration even with increasing task complexity.

Evaluation Highlights

AREX-2 demonstrates strong performance across various benchmarks, particularly in agentic reasoning and engineering tasks. On the MLE-Lite benchmark, it achieves 81.8%, outperforming many larger open-weight models. In general agentic reasoning, it scores 92.2% on GAIA and 93.8% on DeepSearchQA, positioning it competitively among frontier and large models despite its smaller parameter count.

Good for

  • Research on long-horizon agents and iterative problem-solving.
  • Machine-learning engineering and algorithmic coding tasks.
  • Tool-augmented deep research applications.