This dataset was built from YouTube videos with manually provided captions in Cantonese. We used SenseVoice to re-transcribe the audio and filtered segments to build a high-quality collection of audio-caption pairs.
Segments where the ASR output is identical to the original caption — likely clean.
Segments where differences are only homophones (同音字) or English words — likely ASR mistakes.
This combination supports… See the full description on the dataset page:
https://huggingface.co/datasets/ming030890/youtube_caption_yue.