Summary. 212,503-chunk multilingual PII NER dataset covering the
four official Swiss languages and English. 84% is real text from the
Apertus pretrain corpora (Swiss court rulings, federal parliament
records, Swiss-filtered web text, Romansh corpus); the remaining 16%
is template- and LLM-generated synthetic prose used to populate cells
where real-text coverage was insufficient. Annotations are machine
generated by three independent open-weights LLMs (Gemma… See the full description on the dataset page:
https://huggingface.co/datasets/joelbarmettler/gheim-ch-pii-212k.