Synthetic multilingual dataset for training PII (Personally Identifiable Information) detection models with token-level entity annotations and coreference resolution.
Dataset Summary
Samples
99,990 (train: 89,991, test: 9,999)
Languages
6 (Dutch, Spanish, German, English, Danish, French)
Countries
20
PII entity types
26
Total entity annotations
814,306 (avg 8.1 per sample)