AuditoryHuM: Auditory Scene Label Generation and Clustering using Human-MLLM Collaboration.
This is the supplementary material used to generate the results in the paper.
The Keras models require the presence of
https://www.kaggle.com/api/v1/models/google/yamnet/tensorFlow2/yamnet/1/download
Download yamnet-tensorflow2-yamnet-v1.tar.gz and extract this model to a directory named yamnet_model.
mkdir yamnet_model && tar -xvzf… See the full description on the dataset page:
https://huggingface.co/datasets/hzhongresearch/auditoryhum_supplementary.