ipfipfipf/Qwen3.5-4B-sdpo-react-rlsd-multitask-arm5

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3.5-4B is a 4.5 billion parameter multimodal large language model developed by Qwen, featuring a unified vision-language foundation and an efficient hybrid architecture. It supports a native context length of 262,144 tokens, extensible up to 1,010,000 tokens, and excels in reasoning, coding, agent tasks, and visual understanding. The model is designed for robust real-world adaptability and global linguistic coverage across 201 languages and dialects.

Loading preview...

Overview

Qwen3.5-4B is a 4.5 billion parameter multimodal large language model from the Qwen family, designed for exceptional utility and performance. It integrates advancements in multimodal learning, architectural efficiency, and reinforcement learning to offer broad capabilities. The model features a native context length of 262,144 tokens, extensible to 1,010,000 tokens, making it suitable for processing ultra-long texts.

Key Capabilities

  • Unified Vision-Language Foundation: Achieves strong performance across reasoning, coding, agent tasks, and visual understanding benchmarks by early fusion training on multimodal tokens.
  • Efficient Hybrid Architecture: Utilizes Gated Delta Networks and sparse Mixture-of-Experts for high-throughput inference with minimal latency.
  • Scalable RL Generalization: Incorporates reinforcement learning scaled across millions of agent environments for robust real-world adaptability.
  • Global Linguistic Coverage: Supports 201 languages and dialects, enabling inclusive worldwide deployment.
  • Multimodal Input: Capable of processing text, image, and video inputs.

Good For

  • Multimodal Applications: Ideal for tasks requiring understanding and generation across text, images, and videos.
  • Agentic Workflows: Excels in tool calling capabilities, recommended for use with frameworks like Qwen-Agent and Qwen Code.
  • Long Context Processing: Suitable for applications requiring extensive context, with native support for 262,144 tokens and extensibility up to 1,010,000 tokens via YaRN scaling.
  • Multilingual Applications: Strong performance across a wide array of languages and dialects.