The training & validation data behind rizzoaiacademy/rizzo-pii-0.3B
โ a PII token-classification model for Italian legal text, covering 22 categories of
personal data including the Italian legal identifiers (codice fiscale, partita IVA,
dati catastali) that no other open PII model handles.
Everything here serves one goal: anonymize legal documents locally before sending them to a
closed LLM (anonymize โ reversible local dictionary โ API โโฆ See the full description on the dataset page:
https://huggingface.co/datasets/rizzoaiacademy/rizzo-pii-it-dataset.