RedHatAI/Mistral-Small-3.2-24B-Instruct-2506

VISIONConcurrent Unit Cost:2Model Size:24BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

RedHatAI/Mistral-Small-3.2-24B-Instruct-2506 is a 24 billion parameter instruction-tuned language model developed by Mistral AI, building upon Mistral-Small-3.1. This model is specifically enhanced for improved instruction following, reduced repetition errors, and more robust function calling capabilities. It maintains strong performance across various benchmarks, including STEM and vision tasks, making it suitable for complex conversational AI and tool-use applications.

Loading preview...

Overview

RedHatAI/Mistral-Small-3.2-24B-Instruct-2506 is an updated version of Mistral-Small-3.1-24B-Instruct-2503, developed by Mistral AI. This 24 billion parameter model focuses on refining core conversational and functional aspects, offering a more reliable and precise experience for developers.

Key Enhancements

  • Instruction Following: Demonstrates significant improvement in accurately interpreting and executing precise instructions, with Wildbench v2 scores increasing from 55.6% to 65.33% and Arena Hard v2 from 19.56% to 43.1%.
  • Repetition Reduction: Effectively mitigates infinite generations and repetitive outputs, reducing such errors by 2x (from 2.11% to 1.29%) on challenging prompts.
  • Robust Function Calling: Features a more resilient function calling template, enhancing its ability to integrate with external tools and APIs.
  • Multimodal Capabilities: Retains strong vision reasoning capabilities, as evidenced by its performance on benchmarks like ChartQA (87.4%) and DocVQA (94.86%).

Performance Highlights

While primarily focused on instruction following and function calling improvements, the model also shows competitive or slightly improved performance in STEM benchmarks, including a notable increase in HumanEval Plus - Pass@5 to 92.90% and MBPP Plus - Pass@5 to 78.33%.

Recommended Usage

This model is optimized for inference with vLLM (version 0.9.1 or higher) and supports transformers with mistral-common >= 1.6.2. It is particularly well-suited for applications requiring precise instruction adherence, complex function/tool orchestration, and multimodal understanding, especially in scenarios where reducing repetitive outputs is critical. A low temperature (e.g., temperature=0.15) and a well-defined system prompt are recommended for optimal results.