KIEFERSA/Sophea-Qwen3.6-v1

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:2Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

KIEFERSA/Sophea-Qwen3.6-v1 is a 35.1 billion parameter sparse MoE transformer model developed by Kiefer SA (Sophea AI Lab, Athens), fine-tuned from Qwen3.6-35B-A3B with a 32768 token context length. This model specializes in Greek and English reasoning tasks, including math, science, and logic, by generating explicit thinking traces. It retains the full multimodal (vision) capabilities of its base model and includes a multi-token-prediction draft head for speculative decoding, making it suitable for applications requiring auditable step-by-step problem-solving in bilingual contexts.

Loading preview...

Sophea-Qwen3.6-v1: Bilingual Reasoning with Explicit Thinking Traces

Sophea-Qwen3.6-v1, developed by Kiefer SA (Sophea AI Lab, Athens), is a 35.1 billion parameter sparse Mixture-of-Experts (MoE) transformer model fine-tuned from Qwen3.6-35B-A3B. Released alongside the paper "Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See" (arXiv:2608.17744), this model is specifically designed for reasoning tasks in both Greek and English, generating explicit step-by-step thinking traces.

Key Capabilities

  • Bilingual Reasoning: Excels in math, science, and logic problems in both Greek and English, producing reasoning traces in the question's language.
  • Explicit Thinking Traces: Designed to output a detailed reasoning process, which can be enabled via enable_thinking=true for auditable problem-solving.
  • Multimodal (Vision) Support: Inherits the full vision stack from its Qwen3.6-35B-A3B base, allowing for image inputs.
  • Speculative Decoding: Ships with a multi-token-prediction (MTP) draft head for enhanced inference throughput without altering output distribution.
  • Minimal Forgetting: Evaluation shows statistically flat performance on general NLU benchmarks in both Greek and English compared to its base model, indicating no significant loss of general ability.

Good For

  • Greek and English Reasoning Assistants: Ideal for applications requiring robust problem-solving in these languages.
  • Auditable Workflows: Suitable for deployments where the reasoning trace itself is crucial for tutoring, verification, or auditing.
  • Educational Tools: Can be used to demonstrate step-by-step solutions in academic subjects.
  • Multimodal Applications: Leverages its retained vision capabilities for tasks involving both text and image inputs.