PLJIANGG05/qwen3guard-gen-4b-danger-intent-opendata-qwensynth-supplement-sampled4000-v9

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 31, 2026License:cc-by-nc-4.0Architecture:Transformer Open Weights Featherless Exclusive Cold

PLJIANGG05/qwen3guard-gen-4b-danger-intent-opendata-qwensynth-supplement-sampled4000-v9 is a 4 billion parameter Qwen3Guard-Gen model developed by PLJIANGG05, fine-tuned for binary classification of user input as 'SAFE' or 'UNSAFE'. This research classifier, with a 32768 token context length, is specifically designed for danger intent detection, utilizing a diverse dataset including open-source, Qwen-generated, and supplemental samples. It is optimized for classifying user inputs in English and Chinese, serving as a specialized safety classification tool rather than a general-purpose assistant.

Loading preview...

Model Overview

This model, PLJIANGG05/qwen3guard-gen-4b-danger-intent-opendata-qwensynth-supplement-sampled4000-v9, is a 4 billion parameter Qwen3Guard-Gen variant developed by PLJIANGG05. It is specifically fine-tuned for binary classification of user input into SAFE or UNSAFE categories. This is a research classifier, not intended as a general-purpose assistant or an independently certified safety system.

Key Capabilities

  • Danger Intent Classification: Specializes in identifying potentially unsafe user inputs.
  • Binary Output: Provides a clear SAFE or UNSAFE prediction, along with a raw margin and risk score.
  • Optimized for Classification: Uses a fixed, explicit template for binary classification, distinct from generic chat pipelines.
  • Multilingual Data: Trained with English and Chinese task and training data, though expanded comparisons are currently English-only.
  • Context Handling: Supports a maximum classification context of 1,024 tokens, with truncation indicators for longer inputs.

Training Data & Lineage

The model was trained on 4,000 sampled records from an 84,287-record candidate pool. This pool included:

  • opendata: Six public-source datasets (e.g., AdvBench, Aegis V2, HarmBench).
  • qwensynth: Reviewed and corrected Qwen-generated samples.
  • supplement: A second five-category AI-generated supplement.

Good For

  • User Input Safety Screening: Ideal for developers needing to classify user prompts for potential danger intent.
  • Research in Safety Classification: Useful for academic or industrial research into fine-tuned safety models.
  • Integration into Safety Pipelines: Can serve as a component within a broader content moderation or safety system.