ANke121/NAS-PO-GRPO-Qwen3-VL-8B

VISIONPricing:Input $0.727 / Output $5.405Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 24, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

ANke121/NAS-PO-GRPO-Qwen3-VL-8B is an 8 billion parameter Qwen3-VL-Instruct model, fine-tuned using GRPO and Native Attention-Strategy Policy Optimization (NAS-PO) for enhanced vision-language capabilities. This model specializes in multimodal reasoning tasks by decomposing native final-layer attention into visual-mass and conditional visual policies. It achieves 68.04 overall mean Accuracy@8 across out-of-domain multimodal reasoning benchmarks, making it suitable for research in advanced vision-language understanding.

Loading preview...

Model Overview

ANke121/NAS-PO-GRPO-Qwen3-VL-8B is an 8 billion parameter Qwen3-VL-Instruct model that has been specifically fine-tuned using a novel approach combining GRPO with Native Attention-Strategy Policy Optimization (NAS-PO). This model is designed to excel in complex multimodal reasoning tasks by implementing a unique attention decomposition strategy.

Key Capabilities and Innovations

  • Native Attention-Strategy Policy Optimization (NAS-PO): This method decomposes the native final-layer attention into distinct visual-mass and conditional visual policies. It utilizes detached visual mass for trajectory-level advantage scaling and employs AAC (Adaptive Advantage Consolidation) to reinforce successful token-level visual routing.
  • Multimodal Reasoning: The model was trained on the ViRL39K dataset, comprising 38,870 multimodal reasoning examples, to enhance its ability to understand and process information from both visual and textual inputs.
  • Performance: Evaluation reports an overall mean Accuracy@8 of 68.04 across eight configurations from seven out-of-domain multimodal reasoning benchmarks. This metric represents the mean response accuracy.

Training Details

The training involved 202 RLVR steps, using an AdamW optimizer with a learning rate of 1e-6. Specific parameters for NAS-AS included a drop quantile of 0.25 and a minimum scale of 0.9, while AAC used a JSD over q/rho with a target loss ratio of 0.3. The training was conducted on 8 NVIDIA H100 80 GB GPUs.

Intended Use

This model is a research checkpoint, primarily intended for advanced research and development in vision-language understanding and multimodal reasoning. Users should verify outputs before deploying in consequential applications.