tuteliq/Mistral-Small-3.1-24B-Instruct-2503
Mistral-Small-3.1-24B-Instruct-2503 is a 24 billion parameter instruction-tuned model developed by Mistral AI, building upon Mistral Small 3. It integrates state-of-the-art vision understanding and extends its long context capabilities up to 128k tokens, while maintaining strong text performance. This model excels in multimodal tasks, advanced reasoning, and agentic capabilities with native function calling, making it suitable for conversational agents and local inference. It is released under an Apache 2.0 License, supporting dozens of languages.
Loading preview...
Model Overview
Mistral-Small-3.1-24B-Instruct-2503 is an instruction-finetuned model from Mistral AI, featuring 24 billion parameters. It significantly enhances its predecessor, Mistral Small 3, by incorporating state-of-the-art vision understanding and expanding its context window to 128k tokens without compromising text performance. This model is designed to be "knowledge-dense" and can be deployed locally, fitting within a single RTX 4090 or a 32GB RAM MacBook when quantized.
Key Capabilities
- Vision: Analyzes images and provides insights based on visual content alongside text.
- Multilingual: Supports dozens of languages, including English, French, German, Japanese, Chinese, and Arabic.
- Agent-Centric: Offers robust agentic capabilities with native function calling and JSON outputting.
- Advanced Reasoning: Delivers strong conversational and reasoning abilities.
- Long Context: Features a 128k context window, enabling understanding of extensive documents.
- Apache 2.0 License: Allows for broad commercial and non-commercial use and modification.
Performance Highlights
In instruction-tuned evaluations, Mistral-Small-3.1-24B-Instruct-2503 demonstrates competitive performance across various benchmarks:
- Text: Achieves 80.62% on MMLU and 88.41% on HumanEval, indicating strong general knowledge and code generation.
- Vision: Scores 64.00% on MMMU and 68.91% on Mathvista, showcasing its advanced visual understanding.
- Multilingual: Attains an average of 71.18% across diverse language groups.
- Long Context: Achieves 93.96% on RULER 32K and 81.20% on RULER 128K, highlighting its proficiency in processing long inputs.
Good For
- Fast-response conversational agents.
- Low-latency function calling applications.
- Local inference for sensitive data or hobbyist projects.
- Programming and mathematical reasoning tasks.
- Understanding and processing long documents.
- Applications requiring visual understanding and analysis.