Here is the SynParaSpeech dataset.
SynParaSpeech is the first automated synthesis framework for constructing large-scale paralinguistic datasets, designed to solve key issues of existing resources (e.g., missing speech, incomplete annotations, poor realism). It generates high-quality data with 6 fine-grained paralinguistic categories (sigh, throat clearing, laugh, pause, tsk, gasp) that match natural conversational distribution, along with millisecond-level timestamps… See the full description on the dataset page:
https://huggingface.co/datasets/shawnpi/SynParaSpeech.