REAL-T is a conversation-centric benchmark for target speaker extraction (TSE) in real meetings and dinner-party recordings. Unlike simulated stacks such as LibriMix, the mixtures are cut from naturally overlapping speech, with enrollment clips taken from non-overlapping regions of the same conversations.
This package contains the DEV, EVAL1, and EVAL2 splits used by the REAL-TSE Challenge (a satellite challenge… See the full description on the dataset page:
https://huggingface.co/datasets/REAL-TSE/REAL-T.