This dataset documents filler word usage among non‑native English‑speaking university students in Hong Kong during career‑oriented interviews.
It consists of one hour of audio recordings, transcriptions, and lexical‑level annotations highlighting sound, conversational, transitional, and emphasis fillers.
Data were collected through stratified sampling across diverse academic levels, processed with noise reduction, and annotated using a dual AI‑assisted and manual verification approach.… See the full description on the dataset page:
https://huggingface.co/datasets/eduhk-compling/Group_F_Project.