FactualVQA (FVQA) is a multimodal Visual Question Answering dataset created for search-augmented training and evaluation. It emphasizes knowledge-intensive questions that require external information beyond the given image. Each entry includes an image, a question, and an answer (optionally accompanied by candidate answers), enabling models to develop and refine on-demand search strategies. Details of dataset… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/FVQA.