Umranz/Qwen3.8-27B-heretic

VISIONConcurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Qwen3.8-27B-heretic is a 27 billion parameter causal language model developed by Qwen, built on the Qwen3.5 architecture. This model is a native vision-language model, capable of understanding images and videos, and is optimized for complex, multi-step agentic tasks with flexible thinking control. It excels in coding, professional work, research, and long-horizon agentic execution, supporting a native context length of 262,144 tokens, extensible up to 1,000,000 tokens.

Loading preview...

Overview

Qwen3.8-27B is the latest and most capable generation in the Qwen open-model family, building upon the Qwen3.5 architecture. This 27 billion parameter model is a native vision-language model, offering comprehensive improvements across various domains.

Key Capabilities

  • Multimodal Understanding: Natively supports image and video understanding, from STEM diagrams to hour-long videos.
  • Enhanced Agentic Execution: Features stronger autonomous planning and improved handling of environment feedback for more reliable end-to-end task completion.
  • Flexible Thinking Control: Includes a default 'thinking mode' with tunable reasoning depth (reasoning_effort) and preserved reasoning context (preserve_thinking).
  • Broad Compatibility: Offers wider support for popular development tools and harnesses.
  • Extended Context Length: Supports a native context length of 262,144 tokens, extensible up to 1,000,000 tokens using techniques like YaRN.

Performance Highlights

Qwen3.8-27B demonstrates significant gains, particularly in agentic coding (e.g., 61.7% on SWE-bench Pro, 79.0% on QwenSWEBench) and agentic multimodal intelligence (e.g., 84.3% on OSWorld-Verified, 64.8% on WebArena-Verified). It also shows strong performance in general multimodal intelligence, achieving 94.6% on MathVision (with CI) and 85.6% on BabyVision (with CI).

When to Use This Model

This model is ideal for applications requiring advanced multimodal understanding, complex agentic task execution, and long-context processing. Its strengths lie in coding, professional tasks, research, and scenarios where flexible reasoning control is beneficial.