Qwen/Qwen3.5-9B

Hugging Face
VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Feb 27, 2026License:apache-2.0Architecture:Transformer1.7K Open Weights Warm

Qwen3.5-9B is a 9 billion parameter multimodal causal language model developed by Qwen, featuring a unified vision-language foundation and an efficient hybrid architecture. It supports a native context length of 262,144 tokens, extensible up to 1,010,000 tokens, and excels in multimodal reasoning, coding, agentic tasks, and visual understanding. The model also offers expanded linguistic coverage to 201 languages and dialects, making it suitable for global deployment and complex, long-horizon applications.

Loading preview...

Qwen3.5-9B: A Multimodal Agent Foundation Model

Qwen3.5-9B is a powerful 9 billion parameter multimodal causal language model developed by Qwen, designed for exceptional utility and performance across various domains. It integrates significant advancements in multimodal learning, architectural efficiency, and scalable reinforcement learning.

Key Capabilities

  • Unified Vision-Language Foundation: Achieves strong performance in reasoning, coding, agentic tasks, and visual understanding through early fusion training on multimodal tokens.
  • Efficient Hybrid Architecture: Utilizes Gated Delta Networks combined with sparse Mixture-of-Experts for high-throughput inference with reduced latency and cost.
  • Scalable RL Generalization: Features reinforcement learning scaled across millions of agent environments, enhancing real-world adaptability.
  • Global Linguistic Coverage: Supports 201 languages and dialects, enabling inclusive worldwide deployment.
  • Ultra-Long Context: Natively handles up to 262,144 tokens, extensible to 1,010,000 tokens with RoPE scaling techniques like YaRN, making it suitable for complex, long-horizon tasks.
  • Agentic Usage: Excels in tool calling capabilities, with recommended integration via Qwen-Agent and Qwen Code.

Good for

  • Developing multimodal applications requiring strong visual and linguistic understanding.
  • Building AI agents for complex tasks, including coding and general problem-solving.
  • Applications demanding high-throughput inference and efficient resource utilization.
  • Global deployments needing extensive multilingual support.
  • Processing and generating content for extremely long documents or conversations.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p