Dataset Card for Audio Keyword Spotting
Dataset Summary
The initial version of this dataset is a subset of MLCommons/ml_spoken_words, which is derived from Common Voice, designed for easier loading. Specifically, the subset consists of ml_spoken_words files filtered by the names and placenames transliterated in Bible translations, as found in trabina. For our initial experiment, we have focused only on English, Spanish, and Indonesian, three languages whose name… See the full description on the dataset page: https://huggingface.co/datasets/sil-ai/audio-keyword-spotting.