Trendyol/Trendyol-Vision-Master
Trendyol-Vision-Master is a 27 billion parameter vision-language model (VLM) developed by Trendyol, fine-tuned on Qwen3.5-27B. It specializes in e-commerce catalog operations, processing product images, titles, and metadata to generate structured decisions or text outputs. This model excels at tasks like category detection, brand verification, product similarity, and content moderation, primarily supporting Turkish and English languages.
Loading preview...
Trendyol-Vision-Master: E-commerce Catalog VLM
Trendyol-Vision-Master is a 27 billion parameter vision-language model (VLM) developed by the Trendyol Data Science Team, built upon the Qwen3.5-27B architecture. It is specifically fine-tuned for e-commerce catalog quality and moderation workflows, understanding product images, titles, and metadata to produce structured decisions or text outputs.
Key Capabilities
- Category Detection: Selects the best matching category from a provided list based on image and title.
- Product Similarity: Determines if two product listings (images and titles) represent the same SKU.
- Brand Detection: Infers brand names from product images, optionally with category context.
- Attribute Extraction: Extracts structured product attributes from images and titles/descriptions.
- Title Generation: Creates clean catalog titles from product images and reference titles.
- Content Safety Classification: Classifies product content for moderation (Forbidden, Fantasy, Safe).
- Multi-modal Support: Processes both image and text inputs, supporting multi-image scenarios.
When to Use This Model
This model is ideal for automating and enhancing e-commerce catalog operations, particularly for critical or complex cases where high accuracy in category detection and detailed content analysis is required. It is optimized for Trendyol's specific domain and primarily supports Turkish, with secondary support for English. For high-traffic, single-GPU workloads, consider its smaller counterpart, Trendyol-Vision-Flash.