DiscussLLM is a synthetic dataset for the "when to speak" setting in multi-party
discussions. Each example contains a scenario, a discussion transcript, and one
Nexus assistant intervention.
The release contains 88,718 generated discussions with the original split:
@article{patel2025discussllm,
title={DiscussLLM: Teaching Large Language Models… See the full description on the dataset page:
https://huggingface.co/datasets/deepsworld/discussllm.