ArliAI/Qwen3.5-4B-RpRMax-v1

VISIONConcurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Apr 28, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Qwen3.5-4B is a 4.5 billion parameter multimodal causal language model developed by Qwen, featuring a unified vision-language foundation and an efficient hybrid architecture. It excels in multimodal reasoning, coding, and agentic tasks, supporting an extensive 201 languages and dialects. The model offers a native context length of 262,144 tokens, extensible up to 1,010,000 tokens, making it suitable for complex, long-horizon applications requiring deep contextual understanding.

Loading preview...

Qwen3.5-4B: A Multimodal Agentic LLM

Qwen3.5-4B is a 4.5 billion parameter multimodal causal language model from Qwen, designed for advanced utility and performance. It integrates a unified vision-language foundation, allowing for early fusion training on multimodal tokens, which enables it to achieve strong performance across reasoning, coding, agentic tasks, and visual understanding benchmarks.

Key Capabilities & Features

  • Unified Vision-Language Foundation: Achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models in multimodal reasoning, coding, and visual understanding.
  • Efficient Hybrid Architecture: Utilizes Gated Delta Networks and sparse Mixture-of-Experts for high-throughput inference with minimal latency.
  • Scalable RL Generalization: Enhanced real-world adaptability through reinforcement learning scaled across millions of agent environments.
  • Global Linguistic Coverage: Supports 201 languages and dialects for inclusive, worldwide deployment.
  • Extended Context Length: Natively handles 262,144 tokens, extensible up to 1,010,000 tokens using YaRN scaling techniques.
  • Agentic Capabilities: Excels in tool calling, recommended for use with Qwen-Agent and Qwen Code for building agent applications.

When to Use This Model

Qwen3.5-4B is ideal for developers building applications that require:

  • Multimodal Understanding: Processing and reasoning over both text and visual inputs, including images and videos.
  • Complex Reasoning & Coding: Strong performance in mathematical problems, coding challenges, and general instruction following.
  • Agentic Workflows: Leveraging its tool-calling capabilities for automated tasks and agent-based systems.
  • Long Context Processing: Handling extensive documents or conversations with its large context window.
  • Multilingual Applications: Deploying solutions globally with broad language support.