The Nonverbal Vocalization Dataset, also known as the Vocal Characterizer, is a human nonverbal vocal sound dataset that was crowdsourced from the South Korean public and consists of 56.7 hours of brief clips from 1419 speakers. The dataset also contains metadata about sex, age, noise level, and speech quality. "Teeth-chattering," "teeth-grinding," "tongue-clicking," "coughing," "yawning," "throat clearing," "sighing," "lip-popping," "lip-smacking," "panting," "crying," "laughing," "sneezing… See the full description on the dataset page:
https://huggingface.co/datasets/wtfkedar/Non_Verbal_audio.