Raw VideoMME benchmark data plus everything needed to run the CARVE
vanilla-agent / evidence-dependence pipeline against it.
videos/ - raw VideoMME source videos (mp4, sourced from YouTube).
subtitles/ - VideoMME subtitle files.
embeddings/10/large/ - precomputed LanguageBind clip embeddings
(clip_duration=10s, retriever_type=large), keyed by video ID, used
by the retrieval stage so it doesn't need to be recomputed.
qa.json - full… See the full description on the dataset page:
https://huggingface.co/datasets/ramaalhamidi/videomme-carve.