Qwen/Qwen3.8-2.4T-A95B
Qwen/Qwen3.8-2.4T-A95B is a causal language model developed by Qwen, featuring 2.4 trillion total parameters with 95 billion activated, and a native context length of 262,144 tokens, extensible up to 1,010,000. This model is optimized for complex, multi-step agentic tasks, demonstrating comprehensive improvements across coding, professional work, and research. It is designed for reliable end-to-end task completion with enhanced autonomous planning and environment feedback handling.
Loading preview...
Qwen3.8-2.4T-A95B Overview
Qwen3.8-2.4T-A95B is the latest and most capable generation in the Qwen open-model family, building upon the Qwen3.5 architecture. This model introduces a "Qwen-Max-class" model for open release, featuring 2.4 trillion total parameters with 95 billion activated, and a native context length of 262,144 tokens, extensible up to 1,010,000. It is designed for complex, multi-step tasks, offering significant advancements in agentic capabilities.
Key Capabilities & Enhancements
- Core Capabilities: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks.
- Agent Execution: Stronger autonomous planning and improved handling of environment feedback for more reliable end-to-end task completion.
- Flexible Thinking Control: Supports
reasoning_effort(xhigh, medium, low) to tune reasoning depth andpreserve_thinkingto retain reasoning context from historical messages. - Long Context: Natively supports 262,144 tokens, extensible up to 1,010,000 tokens, enabling processing of extensive inputs.
Performance Highlights
Qwen3.8-2.4T-A95B demonstrates strong performance across various benchmarks, particularly in agentic tasks. For instance, it achieves:
- Coding Agent: 86.6 on Terminal Bench 2.1, 67.7 on SWE-bench Pro, and 93.0 on PaperBench.
- General Agent: 74.8 on CoWorkBench and 70.2 on SkillsBench.
- General Capabilities: 92.6 on GPQA Diamond and 82.8 on IFBench.
Usage Considerations
This model is text-only and requires "thinking mode" for all interactions, with reasoning automatically prepended to outputs. It is compatible with popular inference frameworks like SGLang, vLLM, and TokenSpeed. For optimal performance, specific sampling parameters and adequate output length allocation (up to 262,144 tokens for reasoning and 131,072 for final responses) are recommended.