This dataset is the FinQA slice of the UDA (Unstructured Document Analysis) benchmark: 8,190 question–answer instances derived from real financial reports, packaged for retrieval-oriented evaluation in RAG pipelines.
UDA is a benchmark suite for Retrieval-Augmented Generation (RAG) over messy, real-world documents (PDF/HTML) where evidence mixes narrative text and tables. The finance portion includes large subsets aligned with… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/uda_fin_qa.