ivrit.ai is a database of Hebrew audio and text content.
audio-base contains the raw, unprocessed sources.
audio-vad contains audio snippets generated by applying Silero VAD (
https://github.com/snakers4/silero-vad) to the base dataset.
audio-transcripts contains transcriptions for each snippet in the audio-vad dataset.
The audio-base dataset contains data from the following sources:
Geekonomy (Podcast,
https://geekonomy.net)
HaCongress (Podcast,
https://hacongress.podbean.com/)
Idan Eretz's… See the full description on the dataset page:
https://huggingface.co/datasets/ivrit-ai/audio-vad.