halilneed/turkish-pii-detection-v01
halilneed/turkish-pii-detection-v01 is a 270 million parameter model, fine-tuned from cagrigungor/pii-guard-turkish-270m, designed for masking Personal Identifiable Information (PII) in Turkish texts. It processes masking policies provided as instructions, allowing for full, whitelist, or blacklist masking of PII before data is sent to LLMs, log storage, or third parties. This model significantly improves PII detection accuracy, achieving a benchmark score of 0.882 on synthetic Turkish banking/ERP messages.
Loading preview...
Turkish PII Detection Model
halilneed/turkish-pii-detection-v01 is a specialized 270 million parameter model focused on masking Personal Identifiable Information (PII) in Turkish text. It is a continuation of the cagrigungor/pii-guard-turkish-270m model, further fine-tuned on 24,000 synthetic examples targeting previously weaker PII categories.
Key Capabilities
- Instruction-based Masking: The model interprets masking policies provided as instructions, enabling flexible control over which PII fields to mask. This includes:
- Full Masking: Masking all PII in the text.
- Whitelist Masking: Masking only specified PII fields (e.g., IBAN, phone number).
- Blacklist Masking: Masking all PII except specified fields (e.g., all sensitive data except name/surname).
- Category-specific Masking: Masking only special categories of PII.
- Improved Accuracy: Achieved a benchmark score of 0.882 (up from 0.740) on the
cagrigungor/turkish-pii-masking-benchmarkdataset, demonstrating enhanced performance in PII detection and masking. - PII Masking Example: Transforms input like "müşteri Ayşe Yılmaz tc 12345678901 tel 0532 111 22 33" to "müşteri [AD] tc [TCKN] tel [TEL]".
Use Cases
This model is ideal for scenarios requiring privacy protection and data anonymization for Turkish text, particularly before data is processed by Large Language Models (LLMs), stored in logs, or shared with third-party services. It helps ensure compliance by preventing sensitive personal data from being exposed. Developers can find usage examples and an evaluation tool at github.com/halilneed/turkish-pii-detection.
Limitations
While significantly improved, the model shows some weaknesses in texts with multiple individuals, very long texts, or PII with suffixes (e.g., "Cem Aslan'ın"). It is recommended to evaluate the model on specific production data and to break down long texts for optimal performance. Users are responsible for KVKK/GDPR compliance.