trohrbaugh/Qwen3.8-27B-heretic-ara

Hugging Face
VISIONConcurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer0.1K Open Weights Featherless Exclusive Warm

Qwen3.8-27B-heretic-ara is a 27 billion parameter causal language model from the Qwen3.8 family, developed by Qwen. This model is a native vision-language model capable of understanding images and videos, and is specifically enhanced for agentic tasks, coding, professional work, and research. It features flexible thinking control and a native context length of 262,144 tokens, extensible up to 1,000,000 tokens.

Loading preview...

Qwen3.8-27B: Advanced Multimodal Agentic AI

Qwen3.8-27B is a 27 billion parameter causal language model from the Qwen3.8 series, representing the latest generation in the Qwen open-model family. Built upon the Qwen3.5 architecture, this model delivers significant improvements across various domains, particularly in agentic capabilities and multimodal understanding.

Key Capabilities

  • Native Vision-Language Understanding: Processes both images and videos, from STEM diagrams to hour-long video content.
  • Enhanced Agentic Execution: Features stronger autonomous planning and improved handling of environmental feedback, leading to more reliable completion of complex, multi-step tasks.
  • Flexible Thinking Control: Offers configurable reasoning depth (reasoning_effort - xhigh, medium, low) and preserves reasoning context across messages (preserve_thinking). Thinking mode is enabled by default.
  • Broad Compatibility: Designed for easy integration with popular harnesses and development tools like Hugging Face Transformers, vLLM, SGLang, and TokenSpeed.
  • Extended Context Length: Natively supports 262,144 tokens, extensible up to 1,000,000 tokens using techniques like YaRN for ultra-long texts.

Performance Highlights

Qwen3.8-27B demonstrates strong performance across various benchmarks, often outperforming previous Qwen versions and comparable models:

  • Coding: Achieves 61.7 on SWE-bench Pro (agentic coding) and 79.0 on QwenSWEBench (software engineering).
  • Agentic Tasks: Scores 70.7 on CoWorkBench (long-horizon office work) and 33.4 on JobBench (professional job tasks).
  • Multimodal Agentic Intelligence: Leads with 84.3 on OSWorld-Verified (computer use) and 64.8 on WebArena-Verified (browser use).
  • General Multimodal Intelligence: Excels in visual math problem solving (94.6 with CI on MathVision) and general visual reasoning (85.6 with CI on BabyVision).

Good for

  • Complex Agentic Workflows: Ideal for applications requiring autonomous planning and execution of multi-step tasks.
  • Multimodal AI Applications: Suitable for scenarios involving image and video analysis, such as document intelligence, visual web development, and scientific chart analysis.
  • Advanced Coding and Software Engineering: Strong performance in agentic coding and repository-level code generation tasks.
  • Long-Context Processing: Beneficial for tasks requiring extensive context, with support for up to 1M tokens.