A harmonized, multi-source corpus for fine-tuning encoder-style PII / NER
models. 25 canonical PII classes, character-level span annotations,
175,881 English examples, 781,052 entity spans.
Published as a single split (train). Hold-out evaluation is expected to be
performed against unrelated external PII corpora rather than against a slice of
this dataset.
from datasets import load_dataset