Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
kws_dataset_ct – Dataset by heimayuan | AlphaNeural AI
You can deploy this model and start earning money today!
heimayuan
/
kws_dataset_ct
like
0
audio-classification
automatic-speech-recognition
other
audio
webdataset
us
audio
speech
keyword-spotting
kws
webdataset
Views
No views yet
Model card
Files and Versions
Community
API
ygyuan/kws_dataset_ct
Keyword-Spotting (KWS) speech dataset, packed as WebDataset tar shards. The input is a Kaldi-style data directory (wav.scp, text, utt2spk, utt2dur, segments), where each utterance is packed as a single tar sample.
Layout
data/
/ metadata.csv audio/
-000.tar
-001.tar ...
Shard counts:
train: 941 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
.wav # raw… See the full description on the dataset page:
https://huggingface.co/datasets/heimayuan/kws_dataset_ct
.