Qwen/Qwen3.8-27B
Qwen3.8-27B is a 27 billion parameter causal language model developed by Qwen, built upon the Qwen3.5 architecture. This native vision-language model excels in coding, professional tasks, research, and long-horizon agentic tasks, offering native image and video understanding. It features flexible thinking control and is designed for reliable, multi-step task completion, supporting a native context length of 262,144 tokens.
Loading preview...
Qwen3.8-27B: Advanced Multimodal Agentic Capabilities
Qwen3.8-27B is the latest and most capable generation in the Qwen open-model family, a 27 billion parameter causal language model with a native vision encoder. Building on the Qwen3.5 architecture, this model delivers significant improvements across various domains, particularly in agentic tasks and multimodal understanding.
Key Capabilities & Enhancements
- Comprehensive Core Capabilities: Substantial gains in coding, professional work, research, and long-horizon agentic tasks.
- Robust Agent Execution: Features stronger autonomous planning and improved handling of environment feedback, leading to more reliable end-to-end task completion.
- Flexible Thinking Control: Operates in a default 'thinking mode' with adjustable reasoning depth (
reasoning_effort- xhigh, medium, low) and preserved reasoning context (preserve_thinking). This can be disabled for direct responses. - Native Vision-Language Understanding: Supports understanding of both images and videos, including STEM diagrams, documents, and hour-scale video content.
- Extended Context Length: Natively supports a context length of 262,144 tokens, extensible up to 1,000,000 tokens using techniques like YaRN for ultra-long texts.
Performance Highlights
Qwen3.8-27B demonstrates strong performance across various benchmarks, often outperforming comparable models:
- Coding: Achieves 61.7 on SWE-bench Pro (agentic coding) and 79.0 on QwenSWEBench (software engineering).
- Agentic Tasks: Scores 70.7 on CoWorkBench (long-horizon office work) and 33.4 on JobBench (professional job tasks).
- Multimodal Agentic Intelligence: Leads with 84.3 on OSWorld-Verified (computer use), 64.8 on WebArena-Verified (browser use), and 81.9 on AndroidWorld (mobile use).
- General Multimodal Intelligence: Excels in visual math problem solving (94.6 with CI on MathVision) and general visual reasoning (85.6 with CI on BabyVision).
Ideal Use Cases
This model is particularly well-suited for developers and researchers focused on:
- Complex Agentic Workflows: Building agents that require robust planning, execution, and multi-step task completion.
- Multimodal Applications: Developing applications that integrate image and video understanding with advanced language processing.
- Advanced Coding & Software Engineering: Tasks involving agentic coding, repo-level code generation, and software engineering problem-solving.
- Long-Context Processing: Scenarios requiring analysis of extensive documents, codebases, or long video content.