ANke121/NAS-PO-DAPO-Qwen3-VL-8B

VISIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 24, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

ANke121/NAS-PO-DAPO-Qwen3-VL-8B is an 8 billion parameter Qwen3-VL-8B-Instruct model fine-tuned using DAPO and Native Attention-Strategy Policy Optimization (NAS-PO) for vision-language tasks. This model incorporates trajectory-level NAS-AS advantage scaling and positive-anchored token-level AAC routing consolidation. It is designed for vision-language applications, achieving a reported 67.94 overall mean Accuracy@8 across various benchmarks.

Loading preview...

Model Overview

ANke121/NAS-PO-DAPO-Qwen3-VL-8B is an 8 billion parameter vision-language model based on the Qwen3-VL-8B-Instruct architecture. It has been specifically fine-tuned using a novel optimization approach combining DAPO (Deep Advantage Policy Optimization) with NAS-PO: Native Attention-Strategy Policy Optimization for Vision-Language Models.

Key Differentiators

  • NAS-PO Integration: Enhances the DAPO optimization backbone by adding trajectory-level NAS-AS advantage scaling and positive-anchored token-level AAC routing consolidation.
  • Specialized Training: Trained on the ViRL39K dataset, comprising 38,870 examples, with specific rollout and batch configurations.
  • Performance: The model reports a 67.94 overall mean Accuracy@8 across eight configurations from seven benchmarks, indicating its capability in vision-language understanding.

Training Details

  • Data: ViRL39K (38,870 examples)
  • Optimization: AdamW bf16 with a learning rate of 1e-6 and weight decay of 0.01.
  • Hardware: Utilized 8 NVIDIA H100 80 GB GPUs.

Usage Considerations

This model is presented as a research checkpoint. Users should verify outputs before using them in consequential applications. Citation metadata will be provided upon public release of the associated paper.