MksShir/Llama-Guard-3-8B-KSTU-Censor

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 17, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

MksShir/Llama-Guard-3-8B-KSTU-Censor is an 8 billion parameter Llama-3.1-based model, fine-tuned by Meta Llama for content safety classification. It functions as an LLM to classify both prompts and responses, indicating safety status and violated content categories based on the MLCommons standardized hazards taxonomy. This model is optimized for content moderation across 8 languages and supports safety for search and code interpreter tool calls, offering improved performance and lower false positive rates compared to previous versions and GPT-4.

Loading preview...

Overview

Llama-Guard-3-8B-KSTU-Censor is an 8 billion parameter model developed by Meta Llama, based on the Llama-3.1 architecture. It is specifically fine-tuned for content safety classification, acting as an LLM to identify and categorize unsafe content in both user prompts and LLM responses. The model aligns with the MLCommons standardized hazards taxonomy, covering 13 hazard categories plus an additional category for Code Interpreter Abuse.

Key Capabilities

  • Multilingual Content Moderation: Supports content safety classification in 8 languages: English, French, German, Hindi, Italian, Portuguese, Spanish, and Thai.
  • Comprehensive Hazard Detection: Classifies content across 14 categories, including violent crimes, sexual content, hate speech, and specialized advice.
  • Tool Call Safety: Optimized to support safety and security for search and code interpreter tool calls, detecting potential abuse.
  • Improved Performance: Demonstrates higher F1 scores and lower false positive rates compared to Llama Guard 2 and GPT-4 across English, multilingual, and tool use evaluations.

Good For

  • LLM Input/Output Safeguarding: Ideal for developers integrating content moderation into their LLM applications to classify prompts and responses.
  • Multilingual Applications: Suitable for global applications requiring content safety across multiple languages.
  • Enhanced System Safety: Recommended for deployment alongside Llama 3.1 to improve overall system safety, though it may increase benign prompt refusals (false positives).