The two archives implement different training protocols and should not be treated as a controlled single-speaker versus multi-speaker ablation. See the code repository for installation paths and configuration details.
The first row is from the paper. The second row is from the released all-speaker checkpoint. Their protocols differ.
This repository mirrors research artifacts released by the authors. No new license is granted by this model card. The Hugging Face archives intentionally exclude
SMPLX_NEUTRAL_2020.npz; obtain SMPL-X directly from the
official source under its terms. Code, BEAT2, pretrained encoders, and other third-party assets remain subject to their respective terms.
1@inproceedings{zhang2025semtalk,
2 title={SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis},
3 author={Zhang, Xiangyue and Li, Jianfang and Zhang, Jiaxu and Dang, Ziqiang and Ren, Jianqiang and Bo, Liefeng and Tu, Zhigang},
4 booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
5 pages={13761--13771},
6 year={2025},
7 doi={10.1109/ICCV51701.2025.01277}
8}