This is the dataset used for the training of bigcode-pii-model (after training on pseudo-labeled data).
It is a concatenation of an early version of bigcode-pii-dataset which had less samples, and pii-for-code
(a dataset with 400 files we annotated in a previous iteration: MORE INFO TO BE ADDED).
Files with AMBIGUOUS and ID were excluded. Each PII subtype was remaped to it supertype.