This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved… See the full description on the dataset page:
https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-demucs-optimized-terminal-l4-20260814.