uncensoredai/Mistral-Small-24B-Instruct-2501

TEXT GENERATIONConcurrent Unit Cost:2Model Size:24BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Feb 2, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Mistral-Small-24B-Instruct-2501 is a 24 billion parameter instruction-tuned language model developed by Mistral AI, featuring a 32k context window. This model excels in agentic capabilities with native function calling and JSON outputting, alongside advanced multilingual reasoning. It is optimized for fast response conversational agents, low latency function calling, and local inference, making it suitable for deployment on single RTX 4090 GPUs or 32GB RAM MacBooks.

Loading preview...

Mistral-Small-24B-Instruct-2501: An Agent-Centric LLM

Mistral-Small-24B-Instruct-2501, developed by Mistral AI, is a 24 billion parameter instruction-tuned model designed to offer state-of-the-art capabilities in the "small" LLM category. It features a substantial 32k context window and is built upon the Mistral-Small-24B-Base-2501 architecture.

Key Capabilities

  • Agent-Centric Design: Offers best-in-class agentic capabilities, including native function calling and JSON outputting, making it highly suitable for automated workflows.
  • Multilingual Support: Supports dozens of languages, such as English, French, German, Spanish, Italian, Chinese, Japanese, and Korean.
  • Advanced Reasoning: Provides strong conversational and reasoning abilities, maintaining adherence to system prompts.
  • Efficient Deployment: Exceptionally "knowledge-dense" and can be deployed locally on hardware like a single RTX 4090 or a 32GB RAM MacBook, even after quantization.
  • Apache 2.0 License: Available under an open license, permitting both commercial and non-commercial use and modification.

Performance Highlights

Benchmarking indicates strong performance across various categories. In human evaluations against models like Gemma-2-27B and Qwen-2.5-32B, Mistral-Small-24B-Instruct-2501 often performs comparably or better. Public benchmarks show competitive scores in reasoning, knowledge, math, coding (e.g., 0.848 on HumanEval pass@1), and instruction following (e.g., 8.35 on MTBench dev).

Good For

  • Fast Response Conversational Agents: Ideal for applications requiring quick and accurate dialogue.
  • Low Latency Function Calling: Optimized for scenarios where rapid execution of functions is critical.
  • Local Inference: Suitable for hobbyists and organizations handling sensitive data that require on-device processing.
  • Subject Matter Expert Fine-tuning: Provides a strong base for specialized fine-tuning to create domain-specific agents.