A benchmark for always-on, streaming video QA agents.
The user watches a long video (52–134 min) with an AI assistant.
The assistant must answer questions in real time, sometimes based on the current frame,
sometimes by recalling earlier content, sometimes by calling external tools,
and sometimes by proactively waiting for a future moment to speak up.
Overview of StreamArena. Top row: (a) source domains, (b) task types, (c) video durations, (d) temporal gaps… See the full description on the dataset page:
https://huggingface.co/datasets/hkuzxc/StreamArena.