PLJIANGG05/qwen35-4b-danger-intent-opendata-qwensynth-79288-v5
PLJIANGG05/qwen35-4b-danger-intent-opendata-qwensynth-79288-v5 is a 4.5 billion parameter Qwen3.5-based model specifically fine-tuned for binary classification of user input as 'SAFE' or 'UNSAFE'. This research classifier, developed by PLJIANGG05, leverages a diverse training pool of 79,288 records, including open-source datasets and Qwen-generated synthetic data. It is designed for text-input safety classification, providing a risk score and truncation indicator, with a maximum classification context of 1,024 tokens.
Loading preview...
Model Overview
This model, PLJIANGG05/qwen35-4b-danger-intent-opendata-qwensynth-79288-v5, is a 4.5 billion parameter Qwen3.5-based classifier designed to categorize user input as either SAFE or UNSAFE. It is a specialized research tool, not a general-purpose assistant or certified safety system, focusing exclusively on text-input safety classification.
Key Capabilities and Features
- Binary Safety Classification: Determines if user input is
SAFEorUNSAFE. - Extensive Training Data: Fine-tuned on a 79,288-record dataset comprising six public-source datasets (e.g., Aegis V2, OR-Bench, AdvBench) and 29,833 corrected Qwen-generated synthetic samples.
- Scoring Output: Provides a binary prediction, raw margin, risk score, threshold, and truncation indicator.
- Context Handling: Processes a maximum classification context of 1,024 tokens, with longer inputs retaining roughly 75% of the beginning and 25% of the end.
- Multilingual Presence: English and Chinese are present in the training data, though expanded comparisons are currently English-only.
Use Cases and Limitations
This model is primarily intended for research and development in content moderation and safety filtering for text-based inputs. It is crucial to note its limitations:
- Not a General-Purpose Assistant: It does not perform general language generation or understanding tasks.
- No Image/Video Safety Claims: Training and evaluation were exclusively on text; no claims are made regarding image or video safety.
- Research Classifier: Users should validate its performance on independent data and maintain human review for critical decisions.
- Context Truncation: Inputs exceeding 1,024 tokens may lose safety-relevant context due to truncation.