This dataset provides a benchmark for evaluating Vision-Language Models (VLMs) in their ability to guide users to adjust an image to better answer a relevant question.
The Directional Guidance dataset focuses on Visual Question Answering (VQA) tasks where a model needs to evaluate visual information sufficiency and guide the user on where to reposition the camera if the image lacks necessary details. This dataset addresses a unique… See the full description on the dataset page:
https://huggingface.co/datasets/LeoLee7/Directional_Guidance.