nabinkhair/genui-vl-2b-merged

VISIONPricing:Input $0.32 / Cached $0.016 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 25, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The nabinkhair/genui-vl-2b-merged is a 2 billion parameter Qwen3-VL model developed by nabinkhair. This instruction-tuned model was finetuned using Unsloth and Huggingface's TRL library, enabling faster training. It is designed for visual language tasks, leveraging its Qwen3-VL architecture to process both visual and textual inputs. The model is suitable for applications requiring efficient visual language understanding and generation.

Loading preview...

Model Overview

The nabinkhair/genui-vl-2b-merged is a 2 billion parameter visual language model, developed by nabinkhair. It is an instruction-tuned variant of the Qwen3-VL architecture, specifically finetuned from unsloth/Qwen3-VL-2B-Instruct-unsloth-bnb-4bit.

Key Characteristics

  • Architecture: Based on the Qwen3-VL family, indicating its capability to handle both visual and linguistic data.
  • Parameter Count: A compact 2 billion parameters, making it efficient for deployment and inference.
  • Training Efficiency: The model was finetuned using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process.
  • Context Length: Supports a context length of 32768 tokens, allowing for processing of substantial input sequences.

Intended Use Cases

This model is well-suited for applications that require:

  • Visual Language Understanding: Tasks involving the interpretation of images combined with text.
  • Efficient Deployment: Its smaller parameter count makes it a good candidate for resource-constrained environments or applications where speed is critical.
  • Instruction Following: As an instruction-tuned model, it is designed to respond effectively to user prompts and instructions in a visual language context.