zrwang1211/SafeAtlas-Guard-8B
SafeAtlas Guard 8B is an 8 billion parameter target-conditioned multimodal safety model developed by zrwang1211, built upon the Qwen3-VL-8B-Instruct backbone. It evaluates image content, image-grounded user requests, and assistant responses, providing one of five ordered safety levels, a continuous risk score, and auxiliary safety-judge predictions. This model is specifically designed for multimodal safety moderation, ordinal risk assessment, and red-teaming.
Loading preview...
SafeAtlas Guard 8B: Multimodal Safety Model
SafeAtlas Guard 8B is an 8 billion parameter multimodal safety model, part of the SafeAtlas-VL framework, designed to assess and quantify safety risks in multimodal interactions. Built on the Qwen3-VL-8B-Instruct backbone, it processes image content, image-grounded user requests, and assistant responses to provide comprehensive safety evaluations.
Key Capabilities
- Target-Conditioned Safety Evaluation: Evaluates safety based on specific targets (image content, user request, assistant response).
- Five-Level Ordinal Safety Labels: Assigns one of five ordered safety levels:
safe core,safe leaning disputed,boundary uncertain,unsafe leaning disputed, andunsafe core. - Continuous Risk Score: Provides a continuous risk score from 0 to 100, mapping the expected ordinal level.
- Auxiliary Safety-Judge Predictions: Offers three auxiliary safety-judge predictions for request and response targets.
- Multimodal Backbone: Utilizes a Qwen3-VL-8B-Instruct backbone, trained in two stages: full-parameter multimodal instruction tuning for structured safety judgments, followed by training of specific prediction heads.
Performance Highlights
The model demonstrates strong performance across various safety benchmarks, achieving an 80.22% average F1 score on 7 multimodal benchmarks (including BeaverTails-V, SPA-VL, VLGuard, HarmImageTest, LLaVAGuard) and an 80.90% overall average F1 score across 11 benchmarks, including text-only safety evaluations like HarmBench and OpenAI Moderation.
Intended Use Cases
SafeAtlas Guard 8B is primarily intended for:
- Multimodal Safety Moderation
- Ordinal Risk Assessment
- Red-teaming and Evaluation
- Safety Alignment Research
It is crucial to note that the model's training data includes sensitive material, and it should be evaluated in its intended deployment setting, not serving as the sole basis for high-impact decisions.