The piiptbrchatml dataset is designed for training and evaluating models for Personal Identifiable Information (PII) masking in Brazilian Portuguese.
It contains conversations where a system is instructed to mask PII from user inputs. The dataset includes the original text, the masked text, and the identified PII entities.
O dataset piiptbrchatml foi criado para treinar e avaliar modelos para mascaramento de Informações Pessoais Identificáveis… See the full description on the dataset page:
https://huggingface.co/datasets/cicero-im/piiptbrchatml.