zrwang1211/SafeAtlas-Guard-4B

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 25, 2026Architecture:Transformer Featherless Exclusive Cold

SafeAtlas Guard 4B is a 4 billion parameter target-conditioned multimodal safety model developed by zrwang1211, built on the Qwen3-VL-4B-Instruct backbone. It evaluates image content, user requests, and assistant responses, providing one of five safety levels, a continuous risk score, and auxiliary safety-judge predictions. This model is specifically designed for comprehensive multimodal safety moderation and risk assessment.

Loading preview...

SafeAtlas Guard 4B: Multimodal Safety Model

SafeAtlas Guard 4B is a 4 billion parameter multimodal safety model from zrwang1211, utilizing the Qwen3-VL-4B-Instruct as its backbone. It is designed to provide nuanced safety evaluations for various inputs, including images, image-grounded user requests, and assistant responses.

Key Capabilities

  • Comprehensive Safety Assessment: Outputs one of five ordered safety levels (e.g., safe core, unsafe core), a continuous risk score from 0 to 100, and three auxiliary safety-judge predictions.
  • Multimodal Input Evaluation: Capable of analyzing both visual and textual content to determine safety.
  • Target-Conditioned Judgments: Learns structured, target-conditioned safety judgments through a two-stage training process, including full-parameter multimodal instruction tuning.
  • Detailed Prediction Fields: Provides granular data such as scalar latent risk values, learned ordinal thresholds, cumulative probabilities, and probability distributions over safety levels and 16 harm categories.
  • Strong Benchmark Performance: Achieves an overall average F1 score of 80.88% across 11 benchmarks, including BeaverTails-V, SPA-VL, VLGuard, HarmImageTest, LLaVAGuard, HarmBench, and OpenAI Moderation.

Intended Use Cases

SafeAtlas Guard 4B is primarily intended for:

  • Multimodal Safety Moderation: Identifying and mitigating unsafe content across image and text modalities.
  • Ordinal Risk Assessment: Providing a graded evaluation of potential risks.
  • Red-Teaming and Evaluation: Assisting in the discovery of model vulnerabilities and assessing safety performance.
  • Safety Alignment Research: Contributing to the development of safer AI systems.

It is crucial to evaluate the model in its intended deployment setting, as predictions can be context- and policy-dependent, and performance may vary across different languages, cultures, and domains. The model should not be the sole basis for high-impact decisions.