Tuwhy/Qwen3.5-9B-OPSA
Tuwhy/Qwen3.5-9B-OPSA is a 9 billion parameter multimodal causal language model, based on the Qwen3.5 architecture, enhanced with On-Policy Self-Adaptation (OPSA) for improved policy learning without a teacher model or reward. It features a 32,768 token context length, extensible up to 1,010,000 tokens, and excels in reasoning, coding, and multimodal understanding, including image and video inputs. The model demonstrates significant performance gains in mathematical reasoning and general agent tasks due to its OPSA training on text questions from DAPO-17k.
Loading preview...
Model Overview
Tuwhy/Qwen3.5-9B-OPSA is a 9 billion parameter multimodal causal language model built upon the Qwen3.5 architecture, distinguished by its integration of On-Policy Self-Adaptation (OPSA). OPSA is a novel training method that enhances the model's policy without relying on a teacher model, reward model, or task-specific rewards, focusing on improving responses with low actor log probabilities. This specific checkpoint was trained using text questions from DAPO-17k, with a focus on non-thinking mode evaluation.
Key Capabilities & Features
- Enhanced Reasoning: OPSA training significantly boosts performance in mathematical reasoning tasks, with notable improvements on AIME24, AIME25, and HMMT25 benchmarks.
- Multimodal Understanding: Inherits Qwen3.5's unified vision-language foundation, supporting image and video inputs, and demonstrating strong performance across various visual understanding benchmarks like MMMU and MathVision.
- Extended Context Length: Natively supports a context length of 262,144 tokens, extensible up to 1,010,000 tokens using RoPE scaling techniques like YaRN.
- Efficient Architecture: Utilizes a hybrid architecture with Gated Delta Networks and sparse Mixture-of-Experts for high-throughput inference.
- Agentic Capabilities: Excels in tool calling and general agent tasks, with specific support for Qwen-Agent and Qwen Code frameworks.
Ideal Use Cases
- Complex Reasoning: Suited for applications requiring advanced mathematical problem-solving and logical deduction.
- Multimodal AI: Excellent for tasks involving both text and visual (image/video) inputs, such as visual question answering, document understanding, and spatial intelligence.
- Agent Development: Highly effective for building AI agents that require robust tool-calling and adaptive decision-making.
- Long-Context Processing: Ideal for scenarios demanding the processing and generation of ultra-long texts, leveraging its extensible context window.