diyanmuhhamed/Qwen-VL-Coder-R1-7B
diyanmuhhamed/Qwen-VL-Coder-R1-7B is a 7 billion parameter multimodal model that integrates vision, coding, and reasoning capabilities. It is built upon a Qwen2.5-VL language backbone, merged with Qwen2.5-Coder-7B-Instruct and DeepSeek-R1-Distill-Qwen-7B via DARE-TIES, and features its original VL vision tower. This model is designed for tasks requiring both visual understanding and code generation, supporting a context length of 32768 tokens.
Loading preview...
Qwen-VL-Coder-R1-7B Overview
This model, diyanmuhhamed/Qwen-VL-Coder-R1-7B, is a 7 billion parameter multimodal large language model that uniquely combines vision, coding, and reasoning functionalities. It leverages a Qwen2.5-VL language backbone as its foundation, integrating capabilities from both Qwen2.5-Coder-7B-Instruct and DeepSeek-R1-Distill-Qwen-7B through a DARE-TIES merge. The model retains its original VL vision tower, enabling robust visual understanding.
Key Capabilities
- Multimodal Integration: Seamlessly processes both visual and textual inputs.
- Coding Proficiency: Enhanced code generation and understanding due to the integration of coding-focused models.
- Reasoning: Designed to perform complex reasoning tasks across modalities.
- Efficient Deployment: Can be run with
llama.cppor LM Studio, with image support viammproj-Qwen-VL-Coder-R1-7B-f16.gguf, fitting within 8GB of VRAM.
Good For
- Applications requiring visual input interpretation combined with code generation.
- Tasks that benefit from a unified model for vision, coding, and general reasoning.
- Developers looking for a 7B parameter multimodal model that is efficient to deploy on consumer hardware.