IFM/K2-Horizon-MoVA-36B-A4B
IFM/K2-Horizon-MoVA-36B-A4B is a 36 billion parameter Mixture-of-Experts (MoE) model from IFM, featuring Mixture-of-Values attention (MoVA) and activating 4 billion parameters per token. It is designed for agentic and reasoning tasks, demonstrating strong performance against larger open-weight and closed models. The model supports an extensive 512K token context length, making it suitable for complex, long-context applications.
Loading preview...
K2-Horizon-MoVA-36B-A4B Overview
IFM's K2-Horizon-MoVA-36B-A4B is a 36 billion parameter Mixture-of-Experts (MoE) model that utilizes Mixture-of-Values attention (MoVA), activating only 4 billion parameters per token. This architecture allows it to achieve frontier-class results on agentic and reasoning benchmarks, often outperforming open-weight dense and MoE models up to 15 times its size, and competing effectively with closed frontier models.
Key Capabilities
- Efficient Performance: Delivers high performance on reasoning and agentic tasks with a significantly smaller active parameter count (4B per token) compared to its total parameter count (36B).
- Extended Context Length: Features a native 512K token context window, enabling processing of very long inputs and complex information.
- Agentic and Reasoning Excellence: Shows strong benchmark results in agentic tool use (tau3-Banking: 26.8%) and agentic terminal use (Terminal-Bench 2.1: 58.6%), indicating proficiency in complex problem-solving and interaction.
- Open Release: IFM plans to release intermediate checkpoints, training data, and training code, promoting transparency and further research.
Good for
- Agentic Applications: Excels in tasks requiring tool use and interaction with environments, such as banking simulations or terminal operations.
- Complex Reasoning: Suitable for applications demanding high-level reasoning, as evidenced by its performance on benchmarks like Humanity's Last Exam and GPQA Diamond.
- Long-Context Processing: Ideal for use cases that involve analyzing or generating very long documents, code, or conversations due to its 512K context window.
- Resource-Efficient Deployment: Its MoE architecture with 4B active parameters per token offers a balance of performance and computational efficiency compared to dense models of similar total parameter count.