DIH evaluates a model's ability to reason about temporally implicit harm in
short videos: scenes that are individually safe but become unsafe when their
temporal sequence reveals harmful intent or outcome.
The benchmark has two configurations:
DIH-T (Temporal/Visual): 6742 silent video clips, ~29.2 GB.
DIH-M (Multimodal/Audio-Visual): 2983 videos with audio, ~8.0 GB.
Every sample is paired with a binary label (safe / unsafe), a major safety… See the full description on the dataset page:
https://huggingface.co/datasets/dih-neurips/DIH.