NaijaFaceVoiceDB is a large scale dataset of 2,656 subjects(of Nigerian Origin) with over 2 million faces and 195 hours of utterances.
The face and voice datasets can be used independently to build deep learning/transformer based unimodal biometric models while they can be combined to build bimodal models.
Furthermore, the voice datasets can be used to build pretrained speech based GenAI models or finetune existing pretrained models.
The publication based on this large scale dataset in IEEE… See the full description on the dataset page:
https://huggingface.co/datasets/aspmirlab/NaijaFaceVoiceDB.