3.14 million scene-level video clips with multi-level captions, camera labels, semantic tags, quality signals, and duplicate groups.
VidaForge-3M is a large-scale video pretraining dataset produced with
VidaForge, an open data pipeline for
building and studying video foundation model pretraining data.
The release contains 3,141,246 annotation-complete clips totaling
6,475.1 hours. Every clip has four… See the full description on the dataset page:
https://huggingface.co/datasets/VidaForge/VidaForge-3M.