SingMoSub is the first dataset featuring motion subtitles specifically designed for singing-driven 3D head motion. It provides temporally aligned, region-level motion annotations along with acoustic descriptions, enabling detailed modeling of expressive head and facial dynamics in singing scenarios.
The dataset contains over 37 hours of synchronized multimodal data, including:
Singing audio corresponding to each motion sequence;
3D facial motion… See the full description on the dataset page:
https://huggingface.co/datasets/ZikaiHuang/SingMoSub.