This is an annotated dataset for Personal Identifiable Information (PII) in code. The target entities are: Names, Usernames, Emails, IP addresses, Keys, Passwords, and IDs.
The annotation process involved 1,399 crowd-workers from 35 countries with Toloka.
It consists of 12,099 samples of
~50 lines of code in 31 programming languages. You can also find a PII detection model that we trained on this dataset at bigcode-pii-model.… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/bigcode-pii-dataset.