tcz/qwen3-vl-8b-box-layouts-sft-plateau-15000a

VISIONPricing:Input $0.727 / Output $5.405Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 20, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The tcz/qwen3-vl-8b-box-layouts-sft-plateau-15000a is an 8 billion parameter Qwen3-VL model developed by tcz. This vision-language model was fine-tuned using Unsloth and Huggingface's TRL library, enabling faster training. It is designed for tasks leveraging its vision capabilities, building upon the Qwen3-VL architecture.

Loading preview...

Model Overview

The tcz/qwen3-vl-8b-box-layouts-sft-plateau-15000a is an 8 billion parameter vision-language model, developed by tcz. It is a fine-tuned variant of the Qwen3-VL architecture, optimized for specific tasks through supervised fine-tuning.

Key Characteristics

  • Base Model: Finetuned from a Qwen3-VL 8B model.
  • Training Efficiency: The model was trained significantly faster, achieving a 2x speedup, by leveraging the Unsloth library in conjunction with Huggingface's TRL library.
  • Developer: Developed by tcz.
  • License: Distributed under the Apache-2.0 license.

Potential Use Cases

This model is suitable for applications requiring a vision-language understanding, particularly those that benefit from the Qwen3-VL base architecture. Its fine-tuning suggests it is tailored for specific visual layout or box-related tasks, as indicated by its name. Developers looking for an efficient, fine-tuned Qwen3-VL model for visual understanding tasks may find this model useful.