Combined dataset for PII (Personally Identifiable Information) detection,
merging the ai4privacy English-only subset with synthetically generated and semantically validated with different LLMs
challenging examples targeting NER failure modes. Class labels had to be consolidated to prevent label fragmentation too.
ai4privacy/open-pii-masking-500k (English subset): 120,533 train / 30,160… See the full description on the dataset page:
https://huggingface.co/datasets/Ari-S-123/pii-detection-english-consolidated.