A synthetic multi-page visual question answering dataset with hard negatives, designed for fine-tuning visual document retrievers like ColFlor and ColPali.
Banana-Merged contains 1,100 training samples and 10,054 images (positive pages + hard negative variants). Each sample pairs a multi-page analytical query with a set of document images that collectively contain the answer, plus one or more hard negative documents that look visually and… See the full description on the dataset page:
https://huggingface.co/datasets/vkehfdl1/banana-merged.