This dataset is the official benchmark for the paper "The Multilingual Curse at the Retrieval Layer: Evidence from Amharic".
It provides a fixed 90/10 train–test split consisting of 68,000 query–passage pairs, specifically designed for evaluating and training dense, late-interaction, learned sparse, and cross-encoder retrieval models for the Amharic language.
GitHub Repository: rasyosef/amharic-neural-ir
Paper: The Multilingual Curse at the… See the full description on the dataset page:
https://huggingface.co/datasets/rasyosef/Amharic-Passage-Retrieval-Dataset-V2.