large-scale, two-speaker, full-duplex
spoken-dialogue corpus built from public podcast feeds. This repository
distributes no audio
File
Rows
Hours
duplexchat_manifest_en.jsonl.gz
15,304,412
282,634
duplexchat_manifest_ja.jsonl.gz
7,329,011
132,723
manifest_counts.json holds the same totals.
Each line is one two-speaker dialogue clip:
{
"language": "en-us",
"rss_url": "
https://.../podcast/rss"… See the full description on the dataset page:
https://huggingface.co/datasets/sarulab-speech/DuplexChat.