Join our community at
https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Previous world's largest open dataset for privacy. Now it is pii-masking-300k
The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs.
The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion… See the full description on the dataset page:
https://huggingface.co/datasets/saad-kw-almutairi/pii-masking-200k.