XiaomiMiMo/MiMo-V2.5

Hugging Face
VISIONConcurrent Unit Cost:4Model Size:311BQuant:FP8Context Size:32kPublished:Apr 27, 2026License:mitArchitecture:Transformer0.4K Open Weights Warm

MiMo-V2.5 by XiaomiMiMo is a native omnimodal model with a Sparse MoE architecture (310B total / 15B activated parameters) and a 1M token context window. It supports text, image, video, and audio understanding through dedicated encoders, including a 729M-param Vision Transformer and a 261M-param Audio Transformer. This model excels in multimodal perception, long-context reasoning, and agentic workflows, leveraging a hybrid attention architecture and Multi-Token Prediction for efficient inference.

Loading preview...

MiMo-V2.5: Omnimodal Agentic Model

MiMo-V2.5, developed by XiaomiMiMo, is a powerful omnimodal model designed for comprehensive understanding across text, image, video, and audio. Built on the MiMo-V2-Flash backbone, it features a Sparse Mixture of Experts (MoE) architecture with 310 billion total parameters (15 billion activated) and an impressive 1 million token context length.

Key Capabilities & Features

  • Native Omnimodal Understanding: Integrates dedicated 729M-parameter Vision Transformer and 261M-parameter Audio Transformer encoders for high-quality perception across all modalities.
  • Hybrid Attention Architecture: Utilizes a hybrid design of Sliding Window Attention (SWA) and Global Attention (GA) to optimize KV-cache storage while maintaining long-context performance.
  • Efficient Inference: Incorporates Multi-Token Prediction (MTP) modules to accelerate inference via speculative decoding.
  • Agentic Performance: Achieves strong results on agentic tasks and multimodal understanding benchmarks through extensive post-training, including SFT, large-scale agentic RL, and Multi-Teacher On-Policy Distillation (MOPD).
  • Robust Training: Pre-trained on approximately 48 trillion tokens using FP8 mixed precision.

Ideal Use Cases

  • Applications requiring advanced multimodal perception and reasoning.
  • Developing AI agents that interact across various data types (text, image, video, audio).
  • Scenarios demanding long-context understanding and processing.
  • Tasks benefiting from efficient inference and robust agentic capabilities.