A retrieval / grounded-QA benchmark over eight SEC 10-K filings
(AAPL, MSFT, NVDA, AMD; FY2023–24). The distinguishing feature is that
every gold label is derived mechanically from the filing's inline XBRL — not
hand-annotated. No financial-domain judgement enters the labels, so the ground
truth is auditable and reproducible: each answer is a specific tagged fact, and
each question's gold evidence is the set of chunks that contain the XBRL… See the full description on the dataset page:
https://huggingface.co/datasets/jrwana/sec-filing-rag-eval.