Diarized and segmented speech dataset derived from i4ds/spc_r.
Each row is a merged speech segment belonging to a single speaker. The source audio and SRT subtitles from i4ds/spc_r were processed with the following pipeline:
Diarization -- pyannote/speaker-diarization-3.1 assigned speaker labels to each SRT segment based on temporal overlap.
Merging -- Consecutive SRT segments from the same speaker were merged when the silence gap between… See the full description on the dataset page:
https://huggingface.co/datasets/i4ds/spc_r_segmented.