This dataset contains Hebrew text paired with IPA phoneme transcriptions.
The data was produced from ivrit-ai/VoxKnesset. Audio was segmented with Silero VAD, transcribed with ivrit-ai/whisper-large-v3-turbo for Hebrew text, and with renikud/whisper-he-ipa for IPA.
Files
voxknesset-whisper-abjad-he-ipa-raw.tsv
Raw extraction output. Includes source filenames and unnormalized fields.