A ready-to-model, tabular dataset for teaching binary classification on a real speech
problem: detecting the Kyrgyz wake word «Акылай» (Akylai) versus everything else.
Each row is one short audio clip already converted into a fixed-length vector of 250
spectral features, so students can go straight to scikit-learn without touching a single
audio library — yet the problem is a genuine, non-toy… See the full description on the dataset page:
https://huggingface.co/datasets/ainest312/kws-dataset.