Mr. TyDi is a multi-lingual benchmark dataset built on TyDi, covering eleven typologically diverse languages. It is designed for monolingual retrieval, specifically to evaluate ranking with learned dense representations.
This dataset stores documents of Mr. TyDi. To access the queries and judgments, please refer to castorini/mr-tydi.
The only configuration here is the language. As all three folds (train, dev and test) share the same… See the full description on the dataset page:
https://huggingface.co/datasets/castorini/mr-tydi-corpus.