YOYO-AI/Qwen3-VL-4B-YOYO-Instruct

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Dec 19, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

YOYO-AI/Qwen3-VL-4B-YOYO-Instruct is an experimental 4 billion parameter multimodal model from YOYO-AI, designed to explore advanced model merging techniques. It aims to align its visual capabilities with pure text model performance while retaining its visual understanding. This model utilizes a unique ASDF merging method with float32 precision and bfloat16 output, supporting a context length of 262,144 tokens.

Loading preview...

YOYO-AI/Qwen3-VL-4B-YOYO-Instruct: Experimental Multimodal Merging

This model, developed by YOYO-AI, is an experimental 4 billion parameter multimodal model focused on exploring novel merging techniques to unify vision and text capabilities. It aims to achieve text performance comparable to pure text models while maintaining its visual understanding.

Key Technical Highlights:

  • Merging Method: Employs an ASDF merging technique, specifically designed to combine text and vision model tensors.
  • Precision: Utilizes float32 for internal computations and bfloat16 for output, ensuring high precision.
  • Context Length: Features an extended context length of 262,144 tokens, allowing for processing of longer inputs.
  • Merging Algorithm: The core innovation lies in its detailed merging process, which involves:
    • Filtering special layers (embedding, lm_head).
    • Converting tensors to float32 and computing delta.
    • Skipping low-rank tensors.
    • Performing Singular Value Decomposition (SVD) on the difference tensor.
    • Automatically selecting rank via knee point detection on singular values.
    • Reconstructing the delta with top-k components.
    • Fusing the cleaned delta back into the text base.

Underlying Assumption & Future Work:

The merging algorithm operates on the assumption that a model's visual capability is primarily concentrated in a few larger singular values within the residual terms. YOYO-AI notes that a systematic evaluation of the model's visual capabilities is still pending, and this work primarily demonstrates the feasibility of the merging technique. The developers encourage further research into unified vision and text model merging methods.