An open-source evaluation dataset for accurate foreground speaker transcription.
The dataset targets mixture conditions where foreground speech remains generally transcribable by speech-to-text systems, while background speech is distinctly perceived as background. It provides around 90 minutes of foreground–background speech mixtures composed of recorded and synthesized foreground speech, along with ground truth foreground speech and corresponding transcripts.… See the full description on the dataset page:
https://huggingface.co/datasets/ai-coustics/dawn_chorus_en.