This dataset was captured using high-definition cameras to record 208 people's lip and speech videos. It was collected in a quiet indoor environment, simulating various lighting conditions, including normal light, strong light, backlight, and dim light, with shooting distances of 0.5m and 1m, primarily 0.5m, accounting for about 90%. It includes both solo and group recordings. The subjects are mainly Mandarin speakers, with ages ranging from 7 to over 60 years old… See the full description on the dataset page:
https://huggingface.co/datasets/DataoceanAI/Lip_movement_Video_Corpus.