NanoBEIR-ps is a Pashto retrieval benchmark inspired by the NanoBEIR benchmark. It is designed for evaluating semantic search, dense retrieval, sparse retrieval, hybrid search, and Retrieval-Augmented Generation (RAG) systems in the Pashto language.
The dataset provides a lightweight benchmark that can be used to compare embedding models and information retrieval systems without requiring a large-scale corpus.
Supported… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/NanoBEIR-ps.