XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B

Hugging Face
VISIONPricing:Input $0.1078 / Cached $0.0862 / Output $0.28Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026License:mitArchitecture:Transformer0.5K Open Weights Featherless Exclusive Warm

MiMo-V2.6-Distill-Qwen-9B is a 9 billion parameter agentic model developed by Xiaomi MiMo, fine-tuned from Qwen3.5-9B with a 32768 token context length. It specializes in agent tasks, including coding, visual coding, and cybersecurity, demonstrating significant performance improvements over its base model across various benchmarks. This model is designed as an SFT checkpoint for research into agentic reinforcement learning and excels in complex problem-solving across multiple domains.

Loading preview...

Overview

MiMo-V2.6-Distill-Qwen-9B is a 9 billion parameter agentic model developed by Xiaomi MiMo. It is built upon the Qwen3.5-9B architecture and has been extensively fine-tuned using MiMo-generated data, focusing on agentic capabilities across several domains. The model demonstrates substantial performance gains over its base model, particularly in complex agent tasks.

Key Capabilities

  • Agentic Task Execution: Excels in general-purpose agent tasks, showing significant improvements on benchmarks like AutomationBench, Terminal Bench, and Toolathlon-Verified.
  • Coding Proficiency: Achieves strong results in coding challenges, outperforming Qwen3.5-9B on SWE Verified, SWE Pro, and internal MiMo Code benchmarks.
  • Cybersecurity: Demonstrates enhanced capabilities in cybersecurity-related tasks, as evidenced by its performance on the MiMo Cyber (mini) benchmark.
  • Visual Coding: Shows improved performance in visual coding scenarios.
  • Reinforcement Learning Foundation: Released as an SFT checkpoint to serve as a starting point for open research in agentic reinforcement learning.

Training Details

The model was trained on a diverse dataset totaling 77.4 billion tokens, with 27.2 billion loss-bearing tokens. This mixture included substantial contributions from code (29.9% token share), general agent tasks (28.5%), visual tasks (27.4%), and cybersecurity (14.2%).

When to Use This Model

This model is particularly well-suited for:

  • Developing and researching agentic systems that require robust performance in coding, general automation, and cybersecurity.
  • Applications demanding a model capable of complex problem-solving and reasoning across multiple domains.
  • Experiments in agentic reinforcement learning, leveraging its strong SFT foundation.