texdata/Qwen3.6-35B-A3B-Slovenian

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 11, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

The texdata/Qwen3.6-35B-A3B-Slovenian model, developed by MediaAtlas, is a 35.1 billion parameter Qwen3.6-35B-A3B variant specifically fine-tuned for Slovenian language fluency, knowledge, and English-Slovenian translation. This merged full model (bf16) integrates continued-pretraining and supervised fine-tuning to enhance performance in Slovenian-specific tasks. It demonstrates improved accuracy on Slovenian-LLM-Eval benchmarks and significant gains in en↔sl BLEU translation scores, making it suitable for applications requiring robust Slovenian language processing and reasoning.

Loading preview...

texdata/Qwen3.6-35B-A3B-Slovenian: Slovenian-Enhanced Qwen3.6-35B-A3B

This model is a 35.1 billion parameter variant of the Qwen/Qwen3.6-35B-A3B architecture, developed by MediaAtlas. It has undergone continued-pretraining (CPT) and supervised fine-tuning (SFT) to achieve enhanced fluency, knowledge, and translation capabilities for the Slovenian language. The model is provided as a merged full model (bf16), incorporating a LoRA adapter, and preserves multi-token-prediction (MTP) tensors for speculative decoding.

Key Capabilities & Performance

  • Slovenian Language Proficiency: Significantly improved performance on Slovenian-LLM-Eval, with an average accuracy increase of +3.1 points (from 0.623 to 0.654) compared to the base model.
  • Enhanced Translation: Demonstrates notable gains in English↔Slovenian (en↔sl) translation, with BLEU scores improving by +2.47 for en→sl and +4.10 for sl→en.
  • Reasoning Model: Designed to function as a reasoning model, with an option to disable 'thinking' for direct answers or translation.
  • MoE Architecture: Utilizes a Mixture-of-Experts (qwen3_5_moe) architecture.

Usage Considerations

Due to a known issue with fp16/bf16 forward passes in current transformers for this architecture, it is recommended to load the model in 4-bit (bitsandbytes nf4) or use the provided GGUF build for optimal performance. The model's weights are bf16, but it was trained and served under nf4 quantization.