FlagRelease/Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS

TEXT GENERATIONConcurrent Unit Cost:3Model Size:48BQuant:FP8Context Size:32kPublished:Jul 6, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS is a 48 billion parameter large language model developed by MoonshotAI, featuring an innovative hybrid linear attention architecture. Optimized for long-context comprehension, multi-turn dialogue, and complex reasoning, it supports an ultra-long context window up to 1 million tokens. This model is designed for high-efficiency inference, reducing KV cache occupancy while maintaining strong performance on benchmarks. It is particularly suited for long document parsing, knowledge question answering, and industrial intelligent conversation services, with native compatibility for Transformers and vLLM frameworks.

Loading preview...

Kimi-Linear-48B-A3B-Instruct-hygon-FlagOS: High-Efficiency Long-Context LLM

Developed by MoonshotAI, Kimi-Linear-48B-A3B-Instruct is a 48 billion parameter large language model built with a unique hybrid linear attention architecture. This design, featuring a 3:1 structural ratio of Kimi Delta Attention and global MLA, significantly reduces KV cache occupancy and enhances inference throughput, making it highly efficient for demanding applications.

Key Capabilities & Features

  • Ultra-Long Context: Supports an impressive context window of up to 1 million tokens, excelling in long-context comprehension.
  • Optimized for Reasoning: Demonstrates strong performance in complex reasoning scenarios and multi-turn dialogues.
  • High Efficiency: The innovative architecture ensures efficient inference while maintaining robust comprehensive capabilities.
  • FlagOS Integration: Released with a FlagOS-Hygon container image for rapid deployment, supporting a "develop once, run anywhere" workflow across diverse AI accelerators.
  • Benchmark Performance: Achieves competitive results on various authoritative benchmarks, including aime, musr_generative, mmlu_pro, gpqa_generative_cot, and livebench_new.
  • Framework Compatibility: Natively compatible with popular frameworks like Transformers and vLLM.

Good For

  • Long Document Parsing: Ideal for processing and understanding extensive textual content.
  • Knowledge Question Answering: Excels in retrieving and synthesizing information from large knowledge bases.
  • Industrial Intelligent Conversation Services: Suitable for building advanced chatbots and conversational AI systems requiring deep context understanding.
  • Efficient Deployment: Designed for quick and consistent deployment across different hardware environments using the FlagOS stack.