ipfipfipf/Qwen3.5-4B-sdpo-react-rlsd-multitask-arm1.1
Qwen3.5-4B is a 4.5 billion parameter multimodal causal language model developed by Qwen, featuring a unified vision-language foundation and an efficient hybrid architecture. It excels in reasoning, coding, agents, and visual understanding benchmarks, supporting a native context length of 262,144 tokens extensible up to 1,010,000. This model is optimized for robust real-world adaptability and global deployment with expanded linguistic coverage across 201 languages and dialects.
Loading preview...
Qwen3.5-4B Model Overview
Qwen3.5-4B is a 4.5 billion parameter multimodal causal language model from Qwen, designed for advanced utility and performance. It integrates a unified vision-language foundation through early fusion training, achieving strong performance across reasoning, coding, agents, and visual understanding benchmarks, often outperforming previous Qwen3-VL models. The model utilizes an efficient hybrid architecture combining Gated Delta Networks with sparse Mixture-of-Experts for high-throughput inference with minimal latency.
Key Capabilities
- Multimodal Learning: Unified processing of vision and language inputs, including image and video understanding.
- Extended Context: Natively supports 262,144 tokens, extensible up to 1,010,000 tokens using YaRN scaling.
- Global Linguistic Coverage: Expanded support for 201 languages and dialects.
- Agentic Usage: Enhanced tool-calling capabilities, recommended for use with Qwen-Agent and Qwen Code.
- Scalable RL Generalization: Improved real-world adaptability through reinforcement learning across diverse environments.
Good for
- Applications requiring multimodal understanding (image, video, text).
- Long-context tasks and document processing.
- Agent-based systems and tool integration.
- Global deployments needing broad language support.
- Reasoning and coding tasks, as demonstrated by competitive benchmark scores.