MTEB-ready packaging of SQuTR, a bilingual spoken-query-to-text retrieval benchmark.
It contains FiQA, HotpotQA, NQ, MedicalRetrieval, DuRetrieval, and T2Retrieval. Each collection includes clean, 20 dB, 10 dB, and 0 dB audio queries.
Query transcripts are kept for provenance but are not exposed to evaluated MTEB models.
SQuTR's generated data is CC BY-SA 4.0. The packaged corpora and qrels keep their original dataset licenses, including NQ's CC BY-NC-SA 3.0… See the full description on the dataset page:
https://huggingface.co/datasets/artist/SQuTR.