147 questions over five S&P-500 10-K filings (563 printed pages, 1,679 chunks),
where every answerable question carries character-offset gold spans into the
canonical markdown — not just a gold string.
The point of this dataset is not its size. It is that it survived an audit
trail instead of a vibe check, and the trail is published with it.
Most synthetic eval sets are generated once and trusted. This one was generated… See the full description on the dataset page:
https://huggingface.co/datasets/ChihebLovesAi/finrag-eval.