This is the improvised conversational subset of Expresso, machine segmented and transcribed with turn metadata preserved.
Models used:
Silero VAD: initial segmentation of audio files
Parakeet TDT 0.6B V2: transcription of each segment
An 80Hz highpass filter was applied to the source audio prior to segmentation to reduce incidence of the recording environment's low frequency noise.
In total, 27.4 hours of audio remains, as 16790 conversational turns. Some… See the full description on the dataset page:
https://huggingface.co/datasets/nytopop/expresso-improv.