YOYO-AI/Qwen3-VL-4B-YOYO-Instruct
YOYO-AI/Qwen3-VL-4B-YOYO-Instruct is an experimental 4 billion parameter multimodal model from YOYO-AI, designed to explore advanced model merging techniques. It aims to align its visual capabilities with pure text model performance while retaining its visual understanding. This model utilizes a unique ASDF merging method with float32 precision and bfloat16 output, supporting a context length of 262,144 tokens.
Loading preview...
YOYO-AI/Qwen3-VL-4B-YOYO-Instruct: Experimental Multimodal Merging
This model, developed by YOYO-AI, is an experimental 4 billion parameter multimodal model focused on exploring novel merging techniques to unify vision and text capabilities. It aims to achieve text performance comparable to pure text models while maintaining its visual understanding.
Key Technical Highlights:
- Merging Method: Employs an
ASDFmerging technique, specifically designed to combine text and vision model tensors. - Precision: Utilizes
float32for internal computations andbfloat16for output, ensuring high precision. - Context Length: Features an extended context length of
262,144tokens, allowing for processing of longer inputs. - Merging Algorithm: The core innovation lies in its detailed merging process, which involves:
- Filtering special layers (embedding, lm_head).
- Converting tensors to
float32and computing delta. - Skipping low-rank tensors.
- Performing Singular Value Decomposition (SVD) on the difference tensor.
- Automatically selecting rank via knee point detection on singular values.
- Reconstructing the delta with top-k components.
- Fusing the cleaned delta back into the text base.
Underlying Assumption & Future Work:
The merging algorithm operates on the assumption that a model's visual capability is primarily concentrated in a few larger singular values within the residual terms. YOYO-AI notes that a systematic evaluation of the model's visual capabilities is still pending, and this work primarily demonstrates the feasibility of the merging technique. The developers encourage further research into unified vision and text model merging methods.