PII Shield is a large-scale, multilingual dataset for training and evaluating Personally Identifiable Information (PII) detection models. Built by Auren Research, it combines real-world documents from diverse domains with high-quality span-level PII annotations produced by fastino/gliner2-privacy-filter-PII-multi— achieving the highest F1 on the SPY benchmark among open-source PII detectors.
The dataset is designed to… See the full description on the dataset page:
https://huggingface.co/datasets/auren-research/pii-shield.