wkinglin/HalluScope-8B
HalluScope-8B by wkinglin is an 8 billion parameter diagnostic multimodal large language model (MLLM) based on Qwen3-VL-8B-Instruct, designed for fine-grained hallucination detection. It identifies and classifies hallucinated spans within MLLM-generated responses into 12 specific types. This model excels at providing detailed, span-level annotations for hallucination diagnosis in multimodal outputs, making it suitable for MLLM evaluation and safety. It was trained on the HalluScope-30K dataset and supports a 32768 token context length.
Loading preview...
HalluScope-8B: Fine-Grained Hallucination Diagnosis
HalluScope-8B, developed by wkinglin, is an 8 billion parameter multimodal large language model (MLLM) specifically engineered for the fine-grained diagnosis of hallucinations in other MLLM outputs. Built upon the Qwen3-VL-8B-Instruct base model, it processes an image and a model-generated response to detect and classify hallucinated text spans.
Key Capabilities
- Span-level Hallucination Detection: Identifies specific segments of text that contain hallucinations.
- 12 Fine-Grained Hallucination Types: Classifies detected hallucinations into a detailed taxonomy, including:
- Perception: Object, OCR, Numerical_Attribute, Color_Attribute, Shape_Attribute, Spatial_Attribute.
- Reasoning: Logical_Error, Calculation_Error, Knowledge_Error, Query_Misunderstanding, Numerical_Relation, Spatial_Relation.
- Annotated Output: Returns results with hallucinated spans wrapped in typed
<hallucination>tags for clear identification. - Multimodal Input: Accepts both image and text inputs for comprehensive analysis.
Training and Usage
The model was trained using the HalluScope-30K dataset to perform span-level hallucination detection and classification. It is the larger, higher-accuracy variant within the HalluScope family, offering robust diagnostic capabilities. For high-throughput inference, it can be served with vLLM via an OpenAI-compatible API.
Ideal Use Cases
- MLLM Evaluation: Assessing the factual accuracy and reliability of multimodal large language models.
- Content Moderation: Identifying and flagging hallucinated information in AI-generated multimodal content.
- Research: Studying and categorizing different types of MLLM hallucinations.