INDICXNLI is similar to existing XNLI dataset in shape/form, but
focusses on Indic language family.
The train (392,702), validation (2,490), and evaluation sets (5,010) of English
XNLI were translated from English into each of the eleven Indic languages. IndicTrans
is a large Transformer-based sequence to sequence model. It is trained on Samanantar
dataset (Ramesh et… See the full description on the dataset page:
https://huggingface.co/datasets/mteb/IndicXnliPairClassification.