Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
kws_dataset_hm – Dataset by bhyuan | AlphaNeural AI
You can deploy this model and start earning money today!
bhyuan
/
kws_dataset_hm
like
0
audio-classification
automatic-speech-recognition
other
audio
webdataset
us
audio
speech
keyword-spotting
kws
webdataset
Views
No views yet
Model card
Files and Versions
Community
API
ygyuan/kws_dataset_hm
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample.
Layout
data/
/ metadata.csv audio/
-000.tar
-001.tar ...
Shard counts:
train: 114 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
.wav # raw… See the full description on the dataset page:
https://huggingface.co/datasets/bhyuan/kws_dataset_hm
.