Qwen/Qwen3.8-2.4T-A95B

Hugging Face
TEXT GENERATIONPricing:Input $1 / Output $4Concurrent Unit Cost:4Model Size:2400BQuant:FP8Context Size:256kTool Calling:SupportedPublished:Aug 8, 2026License:otherArchitecture:Transformer1.3K Warm

Qwen/Qwen3.8-2.4T-A95B is a causal language model developed by Qwen, featuring 2.4 trillion total parameters with 95 billion activated, and a native context length of 262,144 tokens, extensible up to 1,010,000. This model is optimized for complex, multi-step agentic tasks, demonstrating comprehensive improvements across coding, professional work, and research. It is designed for reliable end-to-end task completion with enhanced autonomous planning and environment feedback handling.

Loading preview...

Qwen3.8-2.4T-A95B Overview

Qwen3.8-2.4T-A95B is the latest and most capable generation in the Qwen open-model family, building upon the Qwen3.5 architecture. This model introduces a "Qwen-Max-class" model for open release, featuring 2.4 trillion total parameters with 95 billion activated, and a native context length of 262,144 tokens, extensible up to 1,010,000. It is designed for complex, multi-step tasks, offering significant advancements in agentic capabilities.

Key Capabilities & Enhancements

  • Core Capabilities: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks.
  • Agent Execution: Stronger autonomous planning and improved handling of environment feedback for more reliable end-to-end task completion.
  • Flexible Thinking Control: Supports reasoning_effort (xhigh, medium, low) to tune reasoning depth and preserve_thinking to retain reasoning context from historical messages.
  • Long Context: Natively supports 262,144 tokens, extensible up to 1,010,000 tokens, enabling processing of extensive inputs.

Performance Highlights

Qwen3.8-2.4T-A95B demonstrates strong performance across various benchmarks, particularly in agentic tasks. For instance, it achieves:

  • Coding Agent: 86.6 on Terminal Bench 2.1, 67.7 on SWE-bench Pro, and 93.0 on PaperBench.
  • General Agent: 74.8 on CoWorkBench and 70.2 on SkillsBench.
  • General Capabilities: 92.6 on GPQA Diamond and 82.8 on IFBench.

Usage Considerations

This model is text-only and requires "thinking mode" for all interactions, with reasoning automatically prepended to outputs. It is compatible with popular inference frameworks like SGLang, vLLM, and TokenSpeed. For optimal performance, specific sampling parameters and adequate output length allocation (up to 262,144 tokens for reasoning and 131,072 for final responses) are recommended.