GutenQA-Propositions consist of the same 100 Public Domain Narrative Books used in GutenQA (the proposed benchmark to the paper LumberChunker: Long-Form Narrative Document Segmentation, and serves as one of the baseline chunking approaches utilized on the LumberChunker paper.
In this version, GutenQA passages are converted to propositions as first described in Dense X Retrieval: What Retrieval Granularity Should We Use.
The dataset is organized into the⦠See the full description on the dataset page:
https://huggingface.co/datasets/LumberChunker/GutenQA_Propositions.