A unified benchmark for vision-language driver alerting — anticipating
driving hazards and emitting graded alerts (SILENT / OBSERVE / ALERT) from
8-frame dashcam clips, each annotated with per-frame safety belief text.
This dataset hosts annotations + experimental results for the VLAlert paper.
Raw videos are not redistributed — see source-dataset links below.
Training/evaluation code is at
AsianPlayer/VLAlert.
Built from 4… See the full description on the dataset page:
https://huggingface.co/datasets/AsianPlayer/VLAlert-Bench.