A micro-sized stratified sample of OpenPII 1.5M,
perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every
label that exists in the parent dataset is represented in proportion.
📖 More information:
www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format
License
100,000
90,000
10,000
19… See the full description on the dataset page:
https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.