A benchmark dataset of 17,083 Croatian parliamentary speech clips from ParlaSpeech-HR v1. Each clip includes audio (WAV) and rich metadata for speaker profiling tasks.
17,083 audio segments
Metadata: speaker info, party affiliation, birth year, gender
Word-level transcriptions and per-word start times (both raw and normalized variants)
Note: ~2,414 audio files referenced in the manifest are currently unavailable in this version… See the full description on the dataset page:
https://huggingface.co/datasets/porupski/ParlaSpeech-HR-benchmark_v1.