QCRI/MemeLens-VLM

VISIONPricing:Input $0.727 / Output $5.405Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 29, 2026License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

QCRI/MemeLens-VLM is an 8 billion parameter unified multilingual, multitask Vision-Language Model (VLM) developed by QCRI for meme understanding. Fine-tuned from Qwen3-VL-8B-Instruct, it utilizes a classify-then-explain strategy on the MemeLens dataset, consolidating 38 public meme datasets across 20 tasks and 9 languages. This model excels at analyzing memes for harm, targets, figurative/pragmatic intent, and affect, demonstrating superior performance over other VLMs and unimodal approaches in these specialized tasks.

Loading preview...

MemeLens-VLM: Multilingual Multitask Meme Understanding

MemeLens-VLM, developed by QCRI, is an 8 billion parameter Vision-Language Model (VLM) specifically designed for comprehensive meme understanding. It is fine-tuned from the Qwen3-VL-8B-Instruct base model using a unique classify-then-explain training strategy.

Key Capabilities & Features

  • Unified Multilingual & Multitask: Consolidates 38 public meme datasets across 20 distinct tasks and 9 languages (AR, BN, DE, EN, ES, HI, RO, RU, ZH).
  • Specialized Meme Analysis: Excels in tasks related to meme understanding, including:
    • Harm: Detecting hateful, harmful, toxic, abusive, and vulgar content.
    • Targets: Identifying targets of misogyny, objectification, shaming, stereotypes, and violence.
    • Figurative/Pragmatic: Analyzing propaganda, metaphor, intention, sarcasm, deepfake, and political content.
    • Affect: Classifying humor, offensiveness, motivational aspects, and sentiment.
  • Superior Performance: Outperforms other multi-modal and uni-modal models, including GPT-4.1 (Zero-Shot) and its base model Qwen3-VL-8B-Instruct, on average across various meme understanding benchmarks, achieving an overall accuracy of 74.1%.
  • Structured Output: Provides classifications in a consistent Label: <label> Explanation: <explanation> format, facilitating structured analysis.

Ideal Use Cases

This model is particularly well-suited for applications requiring in-depth, nuanced analysis of memes, especially in a multilingual context. It can be used for content moderation, social media monitoring, and research into online cultural phenomena where understanding the complex interplay of image and text in memes is crucial.