This repo contains scripts and alignment data to create a dataset build further upon librivoxDeEn such that it contains (German audio, German transcription, English audio, English transcription) quadruplets and can be used for Speech-to-Speech translation research. Because of this, the alignments are released under the same Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License
These alignments were collected by downloading the English audiobooks… See the full description on the dataset page:
https://huggingface.co/datasets/PedroDKE/LibriS2S.