This dataset contains word-level timestamp information for Vietnamese songs, specifically pre-chunked into segments up to 30 seconds for use in training or fine-tuning speech recognition (ASR) systems like Whisper.
The song_dataset_chunked provides high-quality Vietnamese song data, properly segmented into optimal ~30-second sequences.
Duration Insights:
Train split (train_chunked.jsonl): ~ 230.62… See the full description on the dataset page:
https://huggingface.co/datasets/sunbv56/song_dataset_chunked.