This dataset is a Nano-style retrieval dataset for HAKARI-bench.
NanoDAPFAM contains 18 Nano retrieval splits derived from DAPFAM. Each split keeps up to 200 eligible queries and up to 10000 corpus documents, with exact duplicate query and document text removed where the generator records that policy.
dataset_id = "hakari-bench/NanoDAPFAM"
split = "NanoDAPFAMAllTitlAbsClmToFullText"
queries =… See the full description on the dataset page:
https://huggingface.co/datasets/hakari-bench/NanoDAPFAM.