diyanmuhhamed/Qwen-VL-Coder-R1-7B

VISIONConcurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:32kPublished:Jul 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

diyanmuhhamed/Qwen-VL-Coder-R1-7B is a 7 billion parameter multimodal model that integrates vision, coding, and reasoning capabilities. It is built upon a Qwen2.5-VL language backbone, merged with Qwen2.5-Coder-7B-Instruct and DeepSeek-R1-Distill-Qwen-7B via DARE-TIES, and features its original VL vision tower. This model is designed for tasks requiring both visual understanding and code generation, supporting a context length of 32768 tokens.

Loading preview...

Qwen-VL-Coder-R1-7B Overview

This model, diyanmuhhamed/Qwen-VL-Coder-R1-7B, is a 7 billion parameter multimodal large language model that uniquely combines vision, coding, and reasoning functionalities. It leverages a Qwen2.5-VL language backbone as its foundation, integrating capabilities from both Qwen2.5-Coder-7B-Instruct and DeepSeek-R1-Distill-Qwen-7B through a DARE-TIES merge. The model retains its original VL vision tower, enabling robust visual understanding.

Key Capabilities

  • Multimodal Integration: Seamlessly processes both visual and textual inputs.
  • Coding Proficiency: Enhanced code generation and understanding due to the integration of coding-focused models.
  • Reasoning: Designed to perform complex reasoning tasks across modalities.
  • Efficient Deployment: Can be run with llama.cpp or LM Studio, with image support via mmproj-Qwen-VL-Coder-R1-7B-f16.gguf, fitting within 8GB of VRAM.

Good For

  • Applications requiring visual input interpretation combined with code generation.
  • Tasks that benefit from a unified model for vision, coding, and general reasoning.
  • Developers looking for a 7B parameter multimodal model that is efficient to deploy on consumer hardware.