moonshotai/Kimi-K3

Hugging Face
VISIONConcurrent Unit Cost:4Model Size:2780BQuant:FP8Context Size:32kPublished:Jun 13, 2026License:kimi-k3Architecture:Transformer7.8K Warm

Kimi K3 by Moonshot AI is a 2.8 trillion-parameter open-weight, native multimodal agentic model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) with a 1-million-token context window. It features native vision capabilities and utilizes a Stable LatentMoE framework, activating 16 out of 896 experts for improved scaling efficiency. This model is designed for frontier intelligence across long-horizon coding, complex knowledge work, and advanced reasoning tasks.

Loading preview...

Kimi K3: A Frontier Multimodal Agentic Model

Kimi K3, developed by Moonshot AI, is a 2.8 trillion-parameter open-weight multimodal agentic model, notable for being the first open 3T-class model. It integrates Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), alongside a Stable LatentMoE framework that significantly enhances scaling efficiency by activating 16 out of 896 experts. This architecture supports a massive 1-million-token context window and native multimodal capabilities, processing text, images, and video within the same model.

Key Capabilities

  • Long-Horizon Coding: Excels in sustained engineering sessions, navigating large code repositories, and orchestrating terminal tools for tasks like GPU kernel optimization, compiler development, and vision-in-the-loop game development.
  • Agentic Knowledge Work: Capable of producing deep research, interactive visualizations, and motion design, leveraging its native multimodal understanding.
  • Native Multimodality & Long Context: Understands and processes text, images, and video, combined with an industry-leading 1-million-token context window.
  • Advanced Quantization: Utilizes MXFP4 weights with MXFP8 activations through quantization-aware training for broad hardware compatibility.

Good For

  • Developers requiring a powerful model for complex, multi-step coding projects and software engineering tasks.
  • Researchers and professionals needing to perform deep knowledge work, including data analysis, visualization, and content creation.
  • Applications demanding long-context understanding and native multimodal processing (text, image, video) for comprehensive agentic behaviors.