This dataset is a Nano-style Brazilian Portuguese retrieval benchmark for
HAKARI-Bench. It is derived
from the retrieval and reranking tasks exposed by
MTEB-BR at commit e17f36ae304e7bdb76202744c5fb56b440050213. The six public
splits are Quati, JurisTCU, BRTaxQAR, FaQuADIR,
MedPTRetrieval, and FaqBacenRetrieval.
queries = load_dataset(dataset_id… See the full description on the dataset page:
https://huggingface.co/datasets/hakari-bench/NanoMTEB-BR.