A small paired tabular dataset showing the same records before and after
a 10-step anonymization pipeline. Useful as a teaching fixture for privacy
courses, a benchmark for anonymization toolkits, and a sanity-check input
for red-team / membership-inference experiments.
Important: the PII in sample_raw.csv is entirely synthetic.
Names follow the pattern Person_001, emails are
person_001@example.com,
phone numbers are 555-00XX, and "national IDs" are… See the full description on the dataset page:
https://huggingface.co/datasets/t22000t/anonymization-before-after.