gghfez/Mistral-Small-24B-Base-2501
Mistral-Small-24B-Base-2501 by Mistral AI Team is a 24 billion parameter base language model designed for state-of-the-art capabilities in the 'small' LLM category. It features a 32k context window and multilingual support across dozens of languages, excelling in advanced reasoning tasks. This model serves as a strong open-source foundation for various commercial and non-commercial applications.
Loading preview...
Mistral-Small-24B-Base-2501 Overview
Mistral-Small-24B-Base-2501 is a 24 billion parameter base Large Language Model developed by the Mistral AI Team. It is positioned as a leading model in the "small" LLM category, offering capabilities comparable to larger models. This release underscores Mistral AI's commitment to open-source contributions, providing a robust foundation for developers.
Key Features
- Multilingual Support: Capable of processing and generating text in dozens of languages, including English, French, German, Spanish, Italian, Chinese, Japanese, and Korean.
- Advanced Reasoning: Demonstrates state-of-the-art performance in conversational and complex reasoning tasks.
- Extended Context Window: Features a 32,000-token context window, allowing for the processing of longer inputs and maintaining conversational coherence.
- Open License: Released under the Apache 2.0 License, enabling broad commercial and non-commercial use and modification.
- Tokenizer: Utilizes a Tekken tokenizer with a 131k vocabulary size.
Performance Highlights
This model achieves strong benchmark results across various academic evaluations:
- MMLU (5-shot): 80.73
- MMLU Pro (5-shot, CoT): 54.37
- GSM8K (5-shot, maj@1): 80.73
- MBPP (pass@1): 69.64
It also shows competitive multilingual MMLU scores, such as 78.03 for French and 77.69 for German, highlighting its robust cross-lingual understanding.