This dataset is the FeTa-QA slice of the UDA (Unstructured Document Analysis) benchmark: 1,023 question–answer instances grounded in Wikipedia tables and accompanying context, packaged for question answering evaluation in RAG and document-analysis pipelines.
UDA (Hui et al., 2024) is a benchmark suite for Retrieval-Augmented Generation (RAG) over real-world documents in PDF and HTML, where evidence mixes narrative text and… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/uda_feta_qa_qa.