This benchmark is prepared from the 2023 Tatoeba Challenge, by extracting the dev and test sets for languages spoken in the Indian Republic.
The code to download and process the data can be found here in the repo: data_prep/original_v1/extract.py
Note: This is not the official version of Tatoeba benchmark. Just a processed mirror for Indian languages, made available in HuggingFace for ease of use.
Language code… See the full description on the dataset page:
https://huggingface.co/datasets/sarvamai/tatoeba-indic.