Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval
VIRA (Vis-IR Aggregation), a large-scale dataset comprising a vast collection of screenshots from diverse sources, carefully curated into captioned and questionanswer formats.
There are three types of data in VIRA: caption data… See the full description on the dataset page:
https://huggingface.co/datasets/marsh123/VIRA.