A curated test set of 500 deduplicated sentences per language for evaluating language models on Northeast Indian languages.
Assamese (asm) - 500 sentences
Garo (grt) - 500 sentences
Khasi (kha) - 500 sentences
Kokborok (trp) - 500 sentences
Meitei (mni) - 500 sentences
Mizo (lus) - 500 sentences
Naga (nag) - 500 sentences
Nyishi (njz) - 500 sentences
Pnar (pbv) -… See the full description on the dataset page:
https://huggingface.co/datasets/MWirelabs/northeast-languages-test-set.