π World's largest open dataset for privacy masking π
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs.
Key facts:
OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to⦠See the full description on the dataset page:
https://huggingface.co/datasets/AdamiTitus/pii-masking-300k.