sovasoft/zora-v1.12
Sovasoft/zora-v1.12 is an 8 billion parameter open language model built on Qwen3-8B, specifically designed for 12 languages of the Balkans and Southeast Europe. It emphasizes honesty, multi-perspective reasoning, and in-language thinking, admitting when it lacks information rather than hallucinating facts. The model excels in tool-calling, RAG integration, and structured refusals, making it suitable for applications requiring reliable, localized responses in the region's diverse linguistic landscape.
Loading preview...
Zora v1.12: An Honest LLM for the Balkans & Southeast Europe
Zora v1.12, developed by Sovasoft, is an 8 billion parameter open language model built on Qwen3-8B, specifically tailored for 12 languages across the Balkans and Southeast Europe. Its core philosophy centers on honesty, multi-perspective reasoning, and in-language thinking, prioritizing accurate refusals over fabricated responses. This version significantly improves upon its predecessors by integrating Retrieval-Augmented Generation (RAG) and enhanced tool-calling capabilities.
Key Capabilities & Improvements
- Honest & Structured Refusals: Extensive "I don't know" (IDK) training (30-40% of SFT data) enables Zora to provide structured, native-language refusals with reasoning, significantly reducing factual hallucination, particularly in the DETAIL axis of BalkanBench.
- Enhanced Tool-Calling: Features a robust tool-calling cascade (RAG โ web_search โ IDK), allowing the model to access external knowledge bases and web search for more accurate and up-to-date information.
- In-Language Thinking: Trained with synthetic thinking traces to reason directly in the target language, leading to better structured responses.
- Increased Context Length: Supports a MAXLEN of 8192 tokens, an 8x increase from v1.11, providing more room for complex reasoning and longer interactions.
- RAG Integration: Seamlessly integrates with a local RAG knowledge base (87,284 chunks from Wikidata, Wikipedia, EU Law, etc.), providing transparent, sourced, and dated answers.
Performance Highlights
Zora v1.12 achieves a score of 85/156 on BalkanBench, an improvement of +4 points over v1.11. Notable gains were seen in SEARCH (+3) and TOOLBASE (+1) axes, demonstrating the effectiveness of its tool-calling and RAG integration. While FACT, LOGIC, and ANALYSIS remain areas for future improvement due to the 8B capacity limit, the model excels in TEACH, REASON, and LONGFORM tasks.
Ideal Use Cases
- Localized Applications: Perfect for chatbots, customer support, and content generation in Serbian, Croatian, Bosnian, Macedonian, Slovenian, Albanian, Montenegrin, Bulgarian, Greek, Turkish, Romanian, and Hungarian.
- Fact-Checking & Information Retrieval: Its honesty and RAG capabilities make it suitable for tasks where factual accuracy and source transparency are critical.
- Educational Tools: Strong performance in teaching and long-form tasks makes it valuable for educational content and interactive learning platforms.