vidore/vdsid_french dataset that we processed for a visual question answering task where answer is a caption.
@misc{faysse2024colpaliefficientdocumentretrieval,
title={ColPali: Efficient Document Retrieval with Vision Language Models},
author={Manuel Faysse and Hugues Sibille and Tony Wu and Bilel Omrani and Gautier Viaud and Céline Hudelot and Pierre Colombo},
year={2024},
eprint={2407.01449},
archivePrefix={arXiv}… See the full description on the dataset page:
https://huggingface.co/datasets/CATIE-AQ/caption-vidore-vdsid_french-clean.