Synthetic multilingual dataset for training PII (Personally Identifiable Information) detection models with token-level entity annotations and coreference resolution.
Dataset Summary
Samples
51,495 (train: 46,345, test: 5,150)
Languages
6 (English, Danish, Dutch, French, Spanish, German)
Countries
20
PII entity types
26
Total entity annotations
397,441 (avg 7.7 per sample)