This dataset contains precomputed audio features designed for use with the openWakeWord library.
Specifically, they are intended to be used as general purpose negative data (that is, data that does not contain the target wake word/phrase) for training custom openWakeWord models.
The individual .npy files in this dataset are not original audio data, but rather are low dimensional audio features produced by a pre-trained speech embedding model from Google.
openWakeWord uses these features as… See the full description on the dataset page:
https://huggingface.co/datasets/binhpham/livekit_wakeword_features.