π Looking for the open multilingual baseline? Start with
ai4privacy/pii-masking-openpii-1.5m
(1.5M samples, 30 languages, open-PII taxonomy).
Part of PII-Masking-3M by Ai4Privacy, the global
(2M base + Asia Pacific) PII-masking corpus.
π More information:
www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
418,580
2,305,517
26
30
37β¦ See the full description on the dataset page:
https://huggingface.co/datasets/ai4privacy/pii-masking-work-pwi-400k.