This dataset is a Nano-style retrieval dataset for HAKARI-bench.
NanoLaw contains 8 Nano retrieval splits derived from MTEB(Law, v1). Each split keeps up to 200 eligible queries and up to 10000 corpus documents, with exact duplicate query and document text removed where the generator records that policy.
queries = load_dataset(dataset_id, "queries"… See the full description on the dataset page:
https://huggingface.co/datasets/hakari-bench/NanoLaw.