QCRI/ProBel-MTL

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 13, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

QCRI/ProBel-MTL is a 7.6 billion parameter bilingual multi-task model, fine-tuned from Qwen2.5-7B-Instruct by QCRI. It is specifically designed for propaganda detection and persuasion-technique analysis, handling five distinct tasks in both Arabic and English. The model excels at binary propaganda detection, coarse and fine-grained technique classification with explanations, and technique-labeled span extraction in various output formats, making it suitable for research in media analysis.

Loading preview...

QCRI/ProBel-MTL: Bilingual Propaganda Analysis Model

QCRI/ProBel-MTL is a specialized 7.6 billion parameter language model developed by QCRI, fine-tuned from Qwen2.5-7B-Instruct. Its core purpose is to perform comprehensive propaganda detection and persuasion-technique analysis across both Arabic and English texts. This model is unique in its ability to handle five distinct tasks within a single fine-tuned instance, making it a versatile tool for researchers in media and communication studies.

Key Capabilities

  • Bilingual Processing: Supports both Arabic and English inputs and generates responses in the respective language.
  • Multi-Task Proficiency: Addresses five specific propaganda analysis tasks:
    • Binary Propaganda Detection: Classifies text as propaganda (true/false) with an explanation.
    • Coarse-Category Classification: Identifies broad propaganda categories with explanations.
    • Fine-Grained Technique Classification: Pinpoints specific propaganda techniques with explanations.
    • Technique-Labeled Span Extraction (Inline): Tags propaganda spans directly within the input sentence using <span type="Technique">...</span>.
    • Technique-Labeled Span Extraction (JSON): Outputs propaganda spans as a JSON list of {"text", "label", "occurrence"} objects.
  • Context Length: Utilizes a substantial context window of 32768 tokens.
  • Performance: Achieves competitive scores on the ProBel dataset, with Arabic performance generally stronger across tasks (e.g., 0.763 macro-F1 for binary detection in Arabic).

Intended Use Cases

This model is primarily built for research in propaganda and persuasion-technique analysis within news and social media texts. It provides detailed outputs, including explanations and span extractions, which can significantly aid in understanding manipulative language. It is designed to support trained human reviewers rather than replace them, especially in sensitive applications like content moderation.