🚧 This dataset is currently under construction. Specifications may change.
A BM25-first retrieval subset derived from sionic-ai/NanoBEIR-th for reproducible NanoBEIR evaluation.
Language: th
BM25 top-k: 100
Number of evaluation splits: 13
This folder contains generated BM25 retrieval results and reproducibility metadata.
corpus: _id, text
queries: _id, text
qrels: query-id, corpus-id, score… See the full description on the dataset page:
https://huggingface.co/datasets/hotchpotch/NanoBEIR-th-with-bm25.