Summary.MAD is a large-scale video–language dataset created during my PhD research at KAUST (Image and Video Understanding Lab).It provides curated annotations and pre-computed feature representations designed to support tasks such as video grounding, action understanding, and multimodal retrieval.
Maintainer: @soldelliContact:
mattia.soldan@example.com
data/
├── metadata.csv # general information… See the full description on the dataset page:
https://huggingface.co/datasets/soldelli/MAD.