A standardized 8,099-utterance evaluation suite for automatic speech recognition and speech-text retrieval across eight indigenous and regional languages of Northeast India, spanning three language families: Austroasiatic (Khasi), Tibeto-Burman (Garo, Mizo, Kokborok, Wancho, Chakma), and Indo-Aryan (Nagamese, Assamese).
Introduced in NE-MultiSpeech: A Multilingual Speech Corpus and ASR Benchmark for Northeast Indian Languages.
Languages and… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/northeastbench-speech.