Multimodal retrieval training set used to fine-tune visual-document retrieval
embedding models on English document pages: the query is text, the document is a
page image, and each mining row ships 1 positive + 7 mined hard negatives.
Hard negatives were mined with Qwen/Qwen3-VL-Embedding-8B. Mining was performed
within each source dataset, and positives sharing the same query within the same
source dataset were excluded from the… See the full description on the dataset page:
https://huggingface.co/datasets/whybe-choi/en-vdr-hn.