This is the dataset presented in my ASRU-2023 paper.
It consists of multiple files:
Keys2Paragraphs.txt (internal name in scripts: yago_wiki.txt):
4.3 million unique words/phrases (English Wikipedia titles or their parts) occurring in 33.8 million English Wikipedia paragraphs.
Keys2Corruptions.txt (internal name in scripts: sub_misspells.txt):
26 million phrase pairs in the corrupted phrase inventory, as recognized by different ASR models
Keys2Related.txt (internal name in scripts:… See the full description on the dataset page:
https://huggingface.co/datasets/bene-ges/wiki-en-asr-adapt.