magicworld7/qwen-404
Qwen3.8-27B is a 27 billion parameter causal language model developed by Qwen, built upon the Qwen3.5 architecture. This model is a native vision-language model capable of understanding images and videos, and features enhanced agentic capabilities for complex, multi-step tasks. It excels in coding, professional work, research, and long-horizon agentic tasks, offering flexible thinking control and a native context length of 262,144 tokens, extensible up to 1,000,000 tokens.
Loading preview...
Model Overview
Qwen3.8-27B is the latest generation in the Qwen open-model family, a 27 billion parameter causal language model with a vision encoder. Building on the Qwen3.5 architecture, this model delivers significant improvements across various domains, including coding, professional work, research, and long-horizon agentic tasks. It is designed as a compact, deployment-friendly dense model with native vision-language understanding, supporting both images and videos.
Key Capabilities
- Vision-Language Understanding: Native support for interpreting images and videos, from STEM diagrams to hour-scale video content.
- Enhanced Agent Execution: Features stronger autonomous planning and improved handling of environment feedback, leading to more reliable completion of complex, multi-step tasks.
- Flexible Thinking Control: Offers configurable reasoning depth (
reasoning_effort) and retains reasoning context from historical messages (preserve_thinking), with a default 'thinking mode' that can be disabled. - Extended Context Length: Natively supports up to 262,144 tokens, with extensibility up to 1,000,000 tokens using RoPE scaling techniques like YaRN.
- Strong Performance: Achieves leading scores in agentic coding benchmarks like SWE-bench Pro (61.7) and DeepSWE 1.1 (42.2), and agentic multimodal intelligence benchmarks such as OSWorld-Verified (84.3) and WebArena-Verified (64.8).
Good For
- Complex Agentic Workflows: Ideal for applications requiring autonomous planning and execution across various domains, including coding and office tasks.
- Multimodal Applications: Suitable for tasks involving the understanding and processing of both text and visual data (images and videos).
- Long-Context Processing: Beneficial for use cases that demand extensive context windows, such as analyzing large documents or long video sequences.
- Software Engineering Tasks: Excels in agentic coding, repo-level code generation, and software engineering benchmarks.