ZeroAgency/Zero-Mistral-Small-24B-Instruct-2501
ZeroAgency/Zero-Mistral-Small-24B-Instruct-2501 is a 24 billion parameter instruction-tuned causal language model developed by ZeroAgency, based on mistralai/Mistral-Small-24B-Instruct-2501. It is primarily adapted and optimized for performance in both Russian and English languages, trained on the Vikhrmodels/GrandMaster-PRO-MAX dataset. This model excels in conversational tasks and demonstrates improved performance over its base model on Russian-language benchmarks, making it suitable for bilingual applications requiring strong instruction following.
Loading preview...
Overview
Zero-Mistral-Small-24B-Instruct-2501 is an enhanced 24 billion parameter instruction-tuned model from ZeroAgency, built upon the mistralai/Mistral-Small-24B-Instruct-2501 architecture. Its primary distinction lies in its adaptation and optimization for both Russian and English languages, achieved through supervised fine-tuning (SFT) on the GrandMaster-PRO-MAX dataset.
Key Capabilities & Performance
- Bilingual Proficiency: Specifically tuned for high performance in both English and Russian, making it a strong candidate for multilingual applications.
- Improved Benchmarks: Demonstrates notable improvements over its base model, Mistral-Small-24B-Instruct-2501, particularly on Russian-language benchmarks like Ru Arena General and Arena-Hard-Ru lm_eval, where it scores 87.43 and 77.5 respectively.
- Instruction Following: Designed for conversational and chat applications, leveraging its instruction-tuned nature.
- Function Calling: Supports advanced function/tool calling capabilities, as demonstrated in the provided usage examples.
When to Use This Model
- Bilingual Applications: Ideal for use cases requiring robust performance in both Russian and English.
- Conversational AI: Well-suited for chatbots, virtual assistants, and other dialogue-based systems.
- Instruction Following Tasks: Excels in scenarios where precise adherence to instructions is critical.
- Tool Use: Recommended for applications that benefit from function calling and integration with external tools.
Limitations
- Running the 16-bit merged version on GPU requires approximately 55-60 GB of GPU RAM, which may be a consideration for deployment.