This is the cross-lingual Indic subset of the SWIM-IR dataset, where the query generated is in the Indo-European language and the passage is in English.
The SWIM-IR dataset is available as CC-BY-SA 4.0. 18 languages (including English) are available in the cross-lingual dataset.
For full details of the dataset, please read our upcoming NAACL 2024 paper and check out our website.
SWIM-IR dataset is a… See the full description on the dataset page:
https://huggingface.co/datasets/nthakur/indic-swim-ir-cross-lingual.