[ACL'25 Oral] UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
This repository contains the dataset introduced in the paper, consisting of two parts: 5k+ multiple-choice question-answering (MCQ) data and 1k+ video clips.