Dataset Card for Aksharantar
Dataset Summary
Bhasha-Abhijnaanam is a language identification test set for native-script as well as Romanized text which spans 22 Indic languages.
Assamese (asm)
Hindi (hin)
Maithili (mai)
Nepali (nep)
Sanskrit (san)
Tamil (tam)
Bengali (ben)
Kannada (kan)
Malayalam (mal)
Oriya (ori)
Santali (sat)
Telugu (tel)
Bodo(brx)
Kashmiri… See the full description on the dataset page:
https://huggingface.co/datasets/ai4bharat/Bhasha-Abhijnaanam.