zrwang1211/SafeAtlas-Guard-2B

VISIONPricing:Input $0.32 / Cached $0.016 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 25, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

SafeAtlas Guard 2B by zrwang1211 is a 2 billion parameter target-conditioned multimodal safety model, built on Qwen3-VL-2B-Instruct, designed to evaluate image content, user requests, and assistant responses for safety. It provides a five-level safety label, a continuous risk score, and auxiliary safety-judge predictions. This model specializes in comprehensive multimodal safety moderation and ordinal risk assessment with a context length of 32768 tokens.

Loading preview...

SafeAtlas Guard 2B: Multimodal Safety Model

SafeAtlas Guard 2B, developed by zrwang1211, is a 2 billion parameter multimodal safety model introduced in the paper "SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models." It is built upon the Qwen3-VL-2B-Instruct backbone and is specifically engineered to assess the safety of image content, image-grounded user requests, and assistant responses.

Key Capabilities

  • Comprehensive Safety Evaluation: Provides a five-level ordered safety label (e.g., safe core, unsafe core) and a continuous risk score from 0 to 100.
  • Multimodal Analysis: Evaluates safety across images, text requests, and text responses, offering auxiliary safety-judge predictions for request and response targets.
  • Detailed Output: Returns various prediction fields including safety_label, risk_score, category, and teacher_predictions.
  • Robust Training: Utilizes a two-stage training process, including full-parameter multimodal instruction tuning and specialized training for ordinal, harm-category, and teacher-simulation heads.
  • Strong Benchmarks: Achieves an overall average F1 score of 80.43% across 11 benchmarks, including BeaverTails-V, SPA-VL, VLGuard, and HarmBench.

Good For

  • Multimodal Safety Moderation: Ideal for automatically identifying and flagging unsafe content in applications involving images and text.
  • Ordinal Risk Assessment: Provides granular risk scores and safety levels, useful for nuanced content moderation policies.
  • Red-Teaming and Evaluation: Can be used to test the safety boundaries of other AI models and systems.
  • Safety Alignment Research: A valuable tool for researchers working on improving AI safety and alignment.