This dataset was generated by an overlap-aware speaker diarization pipeline.
audio.wav — full audio (mono, 16 kHz) if present
audio.rttm — diarization RTTM if present
summary.csv — per-segment CSV of transcripts if present
speakers/ — cleaned segment WAVs + JSON metadata
from datasets import load_dataset
ds =… See the full description on the dataset page:
https://huggingface.co/datasets/pushthetempo/diarization-52oVLsFJw4s.