KRLabsOrg/acl-verbatim-spans
is a dataset for query-conditioned extractive evidence selection over papers from the
ACL Anthology.
The release combines:
a gold test benchmark with manual span annotations
a larger silver training set produced from synthetic questions, retrieval, and LLM-based
span annotation
an encoder-ready config for training token-classification models directly
The underlying document collection is
KRLabsOrg/acl-anthology-md.… See the full description on the dataset page:
https://huggingface.co/datasets/KRLabsOrg/acl-verbatim-spans.