zrwang1211/SafeAtlas-Guard-4B
SafeAtlas Guard 4B is a 4 billion parameter target-conditioned multimodal safety model developed by zrwang1211, built on the Qwen3-VL-4B-Instruct backbone. It evaluates image content, user requests, and assistant responses, providing one of five safety levels, a continuous risk score, and auxiliary safety-judge predictions. This model is specifically designed for comprehensive multimodal safety moderation and risk assessment.
Loading preview...
SafeAtlas Guard 4B: Multimodal Safety Model
SafeAtlas Guard 4B is a 4 billion parameter multimodal safety model from zrwang1211, utilizing the Qwen3-VL-4B-Instruct as its backbone. It is designed to provide nuanced safety evaluations for various inputs, including images, image-grounded user requests, and assistant responses.
Key Capabilities
- Comprehensive Safety Assessment: Outputs one of five ordered safety levels (e.g.,
safe core,unsafe core), a continuous risk score from 0 to 100, and three auxiliary safety-judge predictions. - Multimodal Input Evaluation: Capable of analyzing both visual and textual content to determine safety.
- Target-Conditioned Judgments: Learns structured, target-conditioned safety judgments through a two-stage training process, including full-parameter multimodal instruction tuning.
- Detailed Prediction Fields: Provides granular data such as scalar latent risk values, learned ordinal thresholds, cumulative probabilities, and probability distributions over safety levels and 16 harm categories.
- Strong Benchmark Performance: Achieves an overall average F1 score of 80.88% across 11 benchmarks, including BeaverTails-V, SPA-VL, VLGuard, HarmImageTest, LLaVAGuard, HarmBench, and OpenAI Moderation.
Intended Use Cases
SafeAtlas Guard 4B is primarily intended for:
- Multimodal Safety Moderation: Identifying and mitigating unsafe content across image and text modalities.
- Ordinal Risk Assessment: Providing a graded evaluation of potential risks.
- Red-Teaming and Evaluation: Assisting in the discovery of model vulnerabilities and assessing safety performance.
- Safety Alignment Research: Contributing to the development of safer AI systems.
It is crucial to evaluate the model in its intended deployment setting, as predictions can be context- and policy-dependent, and performance may vary across different languages, cultures, and domains. The model should not be the sole basis for high-impact decisions.