WildVid-LIP is a large-scale, open-source dataset mapping over 100,000 curated temporal segments from unconstrained, real-world YouTube videos. It provides precise timestamp anchors optimized for training Visual Speech Recognition (VSR / Lip-Reading), audio-visual synchronization, and multimodal self-supervised models.
Instead of distributing heavy, monolithic video files—which introduces platform friction… See the full description on the dataset page:
https://huggingface.co/datasets/Rizul2159/WildVid-LIP.