This dataset is a filtered, re-packaged subset of OpenHumanVid, curated for training camera-controlled, audio-conditioned talking-head video generation models. Each clip is a short talking-head segment where audio, camera trajectory, and portrait are aligned frame-by-frame.
Camera trajectories come from MegaSaM (DepthAnything → UniDepth → DROID-SLAM, whole-person masked) and are metric — in metres… See the full description on the dataset page:
https://huggingface.co/datasets/Haosonnn/OpenHumanVid-Talking.