Fully synthetic two-speaker conversations where the assistant's turn-taking
behaviour is conditioned on a spoken instruction, with construction-time
ground-truth timestamps. Voices are synthesized with
Kokoro-82M (Apache-2.0);
transcripts are synthetic and contain no personal data.
One row is one VARIANT. Rows sharing a base_conv_id share a… See the full description on the dataset page:
https://huggingface.co/datasets/zetianli/FD_data_v2.