Evidence-based benchmark for Kazakh information retrieval — the independent proof base
for the Kazakh Stemmer.
Corpus: 8,370 passages from Kazakh Wikipedia
Queries: 300 queries × 3 categories (natural / inflected / vocabulary-gap)
Format: BEIR-compatible — three subsets: corpus, queries, qrels
Browse the data: use the subset switcher at the top of the Data Studio viewer to
move between corpus (Kazakh passages), queries (the 300 questions), and qrels… See the full description on the dataset page:
https://huggingface.co/datasets/Tim2190/kaz-rag-search-benchmark.