Text corpus for DuplexGen: Adaptive Synthesis of Human–AI Turn-Taking
Dialogues.
This dataset contains DuplexGen-generated dialogues and our own human
turn-taking slot annotations, used to train and calibrate models that
predict when a listener should take the floor, backchannel, or stay silent
during spoken conversation.
A companion dataset, DuplexGen/duplexgen-spoken,
provides a spoken-audio rendering of the generated dialogues (via
Chatterbox TTS). The… See the full description on the dataset page:
https://huggingface.co/datasets/DuplexGen/duplexgen-corpus.