WebTalk-Synthetic is a dataset of synthetic in-the-wild co-speech facial
motion: ~12.7 k short clips, each pairing conversational speech with a
generated 3D facial-motion track in FLAME
coefficient space. The facial motion is produced by an audio-driven face model
from filtered in-the-wild talking audio — it is synthesized, not
motion-captured. The dataset was created for and used by
ViBES (CVPR 2026) to give the face expert
broad in-the-wild coverage.
⚠️ Research… See the full description on the dataset page:
https://huggingface.co/datasets/JuzeZhang/WebTalk-Synthetic.