VoxConverse is an audio-visual diarisation dataset consisting of multispeaker clips of human speech, extracted from YouTube videos. Updates and additional information about the dataset can be found on the dataset website.
Note: This dataset has been preprocessed using diarizers. It makes the dataset compatible with diarizers to fine-tune pyannote segmentation models.
from datasets import load_dataset
ds =… See the full description on the dataset page:
https://huggingface.co/datasets/0x3/voxconverse.