This dataset is the FETA-QA (FeTaQA-style) slice of the UDA (Unstructured Document Analysis) benchmark: 1,023 question–answer instances over Wikipedia-sourced tables, packaged for retrieval-oriented evaluation in RAG pipelines.
UDA (Hui et al., NeurIPS 2024) is a benchmark suite for Retrieval-Augmented Generation (RAG) over messy, real-world documents (PDF/HTML) where evidence mixes narrative text and structured content. In the… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/uda_feta_qa.