VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents
Paper | GitHub
VisBrowse-Bench is a benchmark for visual-native search. It contains 169 VQA instances covering multiple domains and evaluates the models' visual reasoning capabilities during the search process through multimodal evidence cross-validation via text-image retrieval and joint reasoning.
The question and answer fields in the dataset are encrypted. To use the data, you… See the full description on the dataset page:
https://huggingface.co/datasets/Zhengbo-Zhang/VisBrowse-Bench.