FINAL-Bench/Ourbox-35B-JGOS

TEXT GENERATIONConcurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 10, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Ourbox-35B-JGOS is a 35.1 billion parameter Mixture-of-Experts (MoE) reasoning model developed by FINAL-Bench / VIDRAFT_LAB, based on the Qwen3.6-35B-A3B backbone. This model is specialized for Korean language tasks through an evolutionary FFN-level merge process, while maintaining strong performance in English graduate-level science reasoning, achieving 86.36% on GPQA Diamond. It features a hybrid linear/full attention architecture and a 262K token context length, making it suitable for complex, long-context reasoning in Korean and English.

Loading preview...

Ourbox-35B-JGOS: Korean-Specialized Reasoning MoE

Ourbox-35B-JGOS is a 35-billion-parameter Mixture-of-Experts (MoE) model developed by FINAL-Bench / VIDRAFT_LAB using their Darwin evolutionary breeding platform. Built on the Qwen3.6-35B-A3B backbone, this model is uniquely specialized for Korean language tasks through an FFN-level evolutionary merge, rather than traditional retraining.

Key Capabilities & Features

  • Korean Specialization: Enhanced reasoning, comprehension, and generation in Korean, including robust handling of scientific/technical text.
  • High-Performance Reasoning: Achieves 86.36% on GPQA Diamond, a graduate-level science benchmark, outperforming its Qwen3.6-35B-A3B backbone (86.0%) and models like GLM-5.1 (86.2%) with only ~3B active parameters.
  • Darwin Evolutionary Merge: Utilizes a novel method of recombining feed-forward (FFN / MoE expert) tensors from specialized donor models, preserving the backbone's architecture and long-context behavior.
  • Hybrid Attention Architecture: Inherits Qwen3.6-35B-A3B's hybrid attention (¾ linear + ¼ full) for efficient KV-cache scaling across its 262K token context length.
  • Multi-Token Prediction (MTP): Includes an MTP head for built-in draft generation, supporting speculative decoding.
  • Multilingual Support: While Korean-specialized, it retains multilingual capabilities for English, Chinese, Japanese, and other languages from its backbone.

When to Use This Model

  • Korean Language Applications: Ideal for tasks requiring advanced reasoning and generation in Korean.
  • Complex Reasoning Tasks: Suitable for graduate-level scientific and technical problem-solving, even in English.
  • Long-Context Processing: Benefits applications needing to process and understand very long documents or conversations up to 262,144 tokens.
  • Efficient Deployment: Achieves high performance with only ~3B active parameters, making it efficient for its capability level.