Winuim/qwen3-vl-8b-cpt-v2
VISIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 24, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
Winuim/qwen3-vl-8b-cpt-v2 is an 8 billion parameter Qwen3-VL vision-language model developed by Winuim. This model was finetuned using Unsloth and Huggingface's TRL library, enabling faster training. It is designed for tasks requiring both visual and linguistic understanding, building upon the capabilities of the Qwen3-VL architecture.
Loading preview...
Model Overview
Winuim/qwen3-vl-8b-cpt-v2 is an 8 billion parameter vision-language model, finetuned by Winuim. It is based on the Qwen3-VL architecture, specifically finetuned from unsloth/qwen3-vl-8b-instruct-unsloth-bnb-4bit.
Key Characteristics
- Architecture: Qwen3-VL, indicating its capability to process and understand both visual and textual inputs.
- Parameter Count: 8 billion parameters, offering a balance between performance and computational efficiency.
- Training Efficiency: The model was finetuned using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process.
- Context Length: Supports a context length of 32768 tokens, allowing for processing longer sequences of text and potentially more complex visual information.
Intended Use Cases
This model is suitable for applications that require a multimodal understanding, combining visual perception with language processing. Potential use cases include:
- Image captioning and visual question answering.
- Multimodal dialogue systems.
- Content generation based on visual prompts.
- Tasks benefiting from efficient finetuning on a Qwen3-VL base.