ZeroAgency/Zero-Mistral-24B
ZeroAgency/Zero-Mistral-24B is a 24 billion parameter text-only language model developed by ZeroAgency.ru, fine-tuned from Mistral-Small-3.1-24B-Instruct-2503. Optimized for Russian and English, it removes the original model's vision features and retains a long context capability of up to 128k tokens. This model demonstrates good mathematical and reasoning skills, making it suitable for tasks requiring logical problem-solving in both languages.
Loading preview...
Zero-Mistral-24B: A Specialized Multilingual LLM
ZeroAgency/Zero-Mistral-24B is a 24 billion parameter language model developed by ZeroAgency.ru, building upon mistralai/Mistral-Small-3.1-24B-Instruct-2503. This version is specifically adapted for Russian and English languages, with its training primarily leveraging the Big Russian Dataset and proprietary data from Shkolkovo.online.
Key Differentiators & Capabilities
- Text-Only Focus: Unlike its base model, Zero-Mistral-24B has had vision features removed, streamlining it for text-based applications.
- Multilingual Optimization: Enhanced performance for both Russian and English, making it a strong candidate for bilingual applications.
- Mathematical and Reasoning Skills: The model exhibits good capabilities in mathematics and general reasoning tasks.
- Extended Context Window: It maintains the original Mistral model's long context handling, supporting up to 128,000 tokens.
- Performance: Achieves a MERA score of
0.623, with notable results on Russian benchmarks like PARus (0.942 Accuracy) and ruMMLU (0.778 Accuracy).
Recommended Use Cases
- Bilingual Applications: Ideal for tasks requiring robust performance in both Russian and English.
- Problem Solving: Suitable for applications involving mathematical problems and logical reasoning.
- Long-Context Understanding: Effective for processing and generating text over extended conversational histories or documents.
- Virtual Assistants: Can be deployed as a helpful, harmless, and honest virtual assistant, with specific system prompts provided for generic, thought-process, and task-oriented interactions.