This dataset is a curated and preprocessed version of HDTF (High-Definition Talking Face), prepared for tasks such as talking-head generation, video captioning, and multimodal avatar synthesis.
clips.zip
Videos split into 81-frame clips, each representing a short temporal unit
audios.zip
Audio embeddings extracted from original… See the full description on the dataset page:
https://huggingface.co/datasets/Guangtian/HDTF.