This dataset contains lossless FLAC audio embedded in Parquet shards. It was
built from 30 recordings using the text and recording-local speaker labels in
transcript_sbpn, with conservative outer-boundary corrections from ordinary
speaker regions in Diarized.
Chunks: 2,244
Audio duration: 3.682 hours
Speech rows represented: 4,225
Timestamp source: transcript_sbpn_diarized
Merge gap: at most 2.0 seconds
Maximum chunk: 40.0 seconds, except a… See the full description on the dataset page:
https://huggingface.co/datasets/Kppwdfgu1/d1dddfff5fc51ab7.