This repository contains the Visual Perception Question-Answering Dataset (VPQA), introduced in MetaCoder: Encoding Visual Meta Information from Specialist Models for Multimodal LLMs.
Begin by downloading the VPQA Dataset from Hugging Face into your working directory. Then, structure your directory to include the COCO val2017 dataset as follows:… See the full description on the dataset page:
https://huggingface.co/datasets/yingyueli/VPQA.