Predicting the future requires listening as well as seeing.
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding.
FutureOmni is the first benchmark designed… See the full description on the dataset page:
https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.