A large-scale video benchmark of populous, crowded, and chaotic Global South urban environments, used to study world models (JEPA) under soft spatial boundaries, extreme agent heterogeneity, persistent occlusion, and rapid social negotiation.
115,687 clips (4–10 s each)
714 long-form source videos across 22 Indian cities
Drive-through, walk-through, and aerial (drone) viewpoints; markets, ghats, junctions, flyovers, beaches, and more
⚠️ This is a… See the full description on the dataset page: https://huggingface.co/datasets/anonymousML123/denseworld-115k.