GitMylo/nsfwvision-v4_qwen3.5-9b-sft
GitMylo/nsfwvision-v4_qwen3.5-9b-sft is a 9 billion parameter vision-language model based on the Qwen3.5 architecture, fine-tuned for NSFW vision tasks. This model integrates multimodal learning with an efficient hybrid architecture, featuring Gated Delta Networks and sparse Mixture-of-Experts for high-throughput inference. It excels in unified vision-language understanding, outperforming previous Qwen3-VL models across reasoning, coding, agent tasks, and visual benchmarks. The model is designed for applications requiring robust visual understanding and multimodal reasoning, supporting a native context length of 262,144 tokens.
Loading preview...
Model Overview
GitMylo/nsfwvision-v4_qwen3.5-9b-sft is a 9 billion parameter vision-language model built upon the Qwen3.5 foundation, specifically re-merged to retain new vision weights. This model leverages a Unified Vision-Language Foundation with early fusion training on multimodal tokens, achieving strong performance across various benchmarks.
Key Capabilities & Features
- Unified Vision-Language Understanding: Excels in reasoning, coding, agent tasks, and visual understanding, surpassing Qwen3-VL models.
- Efficient Hybrid Architecture: Incorporates Gated Delta Networks and sparse Mixture-of-Experts for optimized inference throughput, minimal latency, and cost efficiency.
- Scalable RL Generalization: Benefits from reinforcement learning scaled across millions of agent environments, enhancing real-world adaptability.
- Extensive Multilingual Support: Expanded to support 201 languages and dialects for global deployment.
- High Context Length: Natively supports a context length of 262,144 tokens, extensible up to 1,010,000 tokens using techniques like YaRN.
- Multimodal Input: Capable of processing text, image, and video inputs, making it versatile for various applications.
Performance Highlights
The Qwen3.5 base model demonstrates strong performance across numerous benchmarks:
- Language Benchmarks: Achieves 82.5 on MMLU-Pro, 88.2 on C-Eval, and 91.5 on IFEval, indicating robust knowledge, reasoning, and instruction-following capabilities.
- Vision-Language Benchmarks: Scores 78.4 on MMMU, 70.1 on MMMU-Pro, and 85.7 on Mathvista (mini), showcasing advanced multimodal reasoning and problem-solving.
- Agentic Capabilities: Shows strong results on BFCL-V4 (66.1) and TAU2-Bench (79.1), highlighting its potential for agent-based applications.
Use Cases
This model is suitable for developers requiring a powerful vision-language model for:
- Multimodal AI applications: Integrating text, image, and video understanding.
- Complex reasoning tasks: Especially those involving visual information.
- Agent development: Leveraging its strong instruction following and tool-calling capabilities.
- Applications requiring long context understanding: Due to its extended context window.