Sidecar dataset for ai-coustics/dawn_chorus_en.
Adds a codec_guess column, applying hypotheses to classify the audio source type(GSM, WhatsApp, Telegram) which is
not present in current dataset, depending on spectral analysis of speech channel on original dataset.
Since audio source type distribution given in the actual dataset(67% GSM, 16.5% WhatsApp, 16.5% Telegram), this
classification is unsupervised, guided by the known prior distribution.… See the full description on the dataset page:
https://huggingface.co/datasets/burak-ozenc/dawn-chorus-codec-labels.