The Polish Interpreting Corpus (PINC) is a hand-verified parallel speech corpus derived from the European Parliament recordings. The corpus was automatically pre-processed and subsequently
manually verified to correct the transcription, word-level speech-to-text alignment and sentence-level interlingual alignment. The audio quality is decent and the annotation is fairly accurate.
The corpus contains a set of 520 recordings of Polish-English speeches… See the full description on the dataset page:
https://huggingface.co/datasets/danijelkorzinek/PINC.