Large Language Models (LLMs) are linked to several issues regarding Personally Identifiable Information (PII). PII
can occur in the training data and can thus be accidentally leaked or extracted with malicious intent, or it can be
inputted in LLM-based technologies by users through their prompts. A viable strategy to limit the LLMs exposure to
PII is to filter input and output data by de-identifying PII, including personal names. This however poses a challenge:… See the full description on the dataset page:
https://huggingface.co/datasets/IIS-NLP/VEIL.