805,000 synthetic, exactly-labeled PII token-classification examples across
~23 languages and 6 scripts, built for training
nym's PII detection models
(e.g. Wismut/nym-pii-multilingual).
JSONL with character-offset spans (offsets index into text as UTF-8 —
compatible with HF fast-tokenizer offset_mapping):
{"text": "Passport Y94316756 issued to Gary Fisher, Ukraine, expires 11/03/2008.",
"entities": [{"start": 9, "end": 18… See the full description on the dataset page:
https://huggingface.co/datasets/Wismut/nym-pii-multilingual-data.