ZeroAgency/Zero-Mistral-24B

TEXT GENERATIONPricing:Input $0.7 / Cached $0.04 / Output $1.16Concurrent Unit Cost:2Model Size:24BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 21, 2025License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

ZeroAgency/Zero-Mistral-24B is a 24 billion parameter text-only language model developed by ZeroAgency.ru, fine-tuned from Mistral-Small-3.1-24B-Instruct-2503. Optimized for Russian and English, it removes the original model's vision features and retains a long context capability of up to 128k tokens. This model demonstrates good mathematical and reasoning skills, making it suitable for tasks requiring logical problem-solving in both languages.

Loading preview...

Zero-Mistral-24B: A Specialized Multilingual LLM

ZeroAgency/Zero-Mistral-24B is a 24 billion parameter language model developed by ZeroAgency.ru, building upon mistralai/Mistral-Small-3.1-24B-Instruct-2503. This version is specifically adapted for Russian and English languages, with its training primarily leveraging the Big Russian Dataset and proprietary data from Shkolkovo.online.

Key Differentiators & Capabilities

  • Text-Only Focus: Unlike its base model, Zero-Mistral-24B has had vision features removed, streamlining it for text-based applications.
  • Multilingual Optimization: Enhanced performance for both Russian and English, making it a strong candidate for bilingual applications.
  • Mathematical and Reasoning Skills: The model exhibits good capabilities in mathematics and general reasoning tasks.
  • Extended Context Window: It maintains the original Mistral model's long context handling, supporting up to 128,000 tokens.
  • Performance: Achieves a MERA score of 0.623, with notable results on Russian benchmarks like PARus (0.942 Accuracy) and ruMMLU (0.778 Accuracy).

Recommended Use Cases

  • Bilingual Applications: Ideal for tasks requiring robust performance in both Russian and English.
  • Problem Solving: Suitable for applications involving mathematical problems and logical reasoning.
  • Long-Context Understanding: Effective for processing and generating text over extended conversational histories or documents.
  • Virtual Assistants: Can be deployed as a helpful, harmless, and honest virtual assistant, with specific system prompts provided for generic, thought-process, and task-oriented interactions.