IFM/K2-Horizon-MoVA-36B-A4B

Hugging Face
TEXT GENERATIONPricing:Input $1 / Output $4Concurrent Unit Cost:3Model Size:36BQuant:FP8Context Size:32kPublished:Sep 1, 2026License:apache-2.0Architecture:Transformer0.3K Open Weights Warm

IFM/K2-Horizon-MoVA-36B-A4B is a 36 billion parameter Mixture-of-Experts (MoE) model from IFM, featuring Mixture-of-Values attention (MoVA) and activating 4 billion parameters per token. It is designed for agentic and reasoning tasks, demonstrating strong performance against larger open-weight and closed models. The model supports an extensive 512K token context length, making it suitable for complex, long-context applications.

Loading preview...

K2-Horizon-MoVA-36B-A4B Overview

IFM's K2-Horizon-MoVA-36B-A4B is a 36 billion parameter Mixture-of-Experts (MoE) model that utilizes Mixture-of-Values attention (MoVA), activating only 4 billion parameters per token. This architecture allows it to achieve frontier-class results on agentic and reasoning benchmarks, often outperforming open-weight dense and MoE models up to 15 times its size, and competing effectively with closed frontier models.

Key Capabilities

  • Efficient Performance: Delivers high performance on reasoning and agentic tasks with a significantly smaller active parameter count (4B per token) compared to its total parameter count (36B).
  • Extended Context Length: Features a native 512K token context window, enabling processing of very long inputs and complex information.
  • Agentic and Reasoning Excellence: Shows strong benchmark results in agentic tool use (tau3-Banking: 26.8%) and agentic terminal use (Terminal-Bench 2.1: 58.6%), indicating proficiency in complex problem-solving and interaction.
  • Open Release: IFM plans to release intermediate checkpoints, training data, and training code, promoting transparency and further research.

Good for

  • Agentic Applications: Excels in tasks requiring tool use and interaction with environments, such as banking simulations or terminal operations.
  • Complex Reasoning: Suitable for applications demanding high-level reasoning, as evidenced by its performance on benchmarks like Humanity's Last Exam and GPQA Diamond.
  • Long-Context Processing: Ideal for use cases that involve analyzing or generating very long documents, code, or conversations due to its 512K context window.
  • Resource-Efficient Deployment: Its MoE architecture with 4B active parameters per token offers a balance of performance and computational efficiency compared to dense models of similar total parameter count.