MMDocRAG: Benchmarking Retrieval-Augmented Multimomal Generation for Document Question Answering
Kuicai Dong*
·
Yujing Chang*
·
Shijie Huang
·
Yasheng Wang
·
Ruiming Tang
·
Yong Liu
📖Paper |🏠Homepage|👉Github
Document Visual Question Answering (DocVQA) faces dual challenges in processing lengthy multimodal documents (text, images, tables) and performing cross-modal reasoning. Current document retrieval-augmented generation… See the full description on the dataset page: https://huggingface.co/datasets/MMDocIR/MMDocRAG.