LequeuISIR/AU-extraction_Qwen2.5-7B-Instruct
LequeuISIR/AU-extraction_Qwen2.5-7B-Instruct is a 7.6 billion parameter Qwen2.5-7B-Instruct model fine-tuned by LequeuISIR for Argumentative Unit (AU) Extraction. This model is specifically optimized for segmenting French opinion texts into argumentative units, making it the best performing model for this task on the GDN-CC dataset. It excels at identifying and extracting non-contiguous text segments that form coherent argumentative units, providing a structured output for downstream analysis.
Loading preview...
Model Overview
LequeuISIR/AU-extraction_Qwen2.5-7B-Instruct is a specialized 7.6 billion parameter Qwen2.5-7B-Instruct model developed by LequeuISIR. It has been fine-tuned on the GDN-CC dataset, focusing on the task of Argumentative Unit Extraction from French opinion texts. This model is noted as the top performer for AU extraction and was used to annotate the GDN-CC-large corpus.
Key Capabilities
- Argumentative Unit Extraction: Segments French opinion texts into argumentative units, which can include solutions, arguments, or simple affirmations.
- Non-Contiguous Segment Handling: Capable of identifying and concatenating non-contiguous text segments that form a single argumentative unit.
- Extractive Task: Strictly copies and extracts text segments exactly as they appear in the original document, preserving capitalization and punctuation.
- Structured Output: Presents extracted argumentative units as a list, facilitating further processing and analysis.
Use Cases
This model is particularly well-suited for:
- Analysis of Democratic Citizen Consultations: Designed to process noisy, multi-topic contributions from public forums into structured argumentative units.
- Opinion Clustering and Political Analysis: Preprocessing data for topic modeling and political analysis by standardizing citizen contributions at a pragmatic level.
- Research on Argumentation Mining: Provides a robust tool for extracting argumentative structures from French text, as demonstrated by its use in annotating the GDN-CC-large dataset.
Technical Details
The model is recommended for use with the vLLM framework for efficient inference. The fine-tuning process and dataset are detailed in the associated research paper, "The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations" by Lequeu et al. (2026).