AuditoryBench++ is a benchmark designed to evaluate auditory commonsense knowledge and reasoning abilities of language models without requiring direct audio input.Humans can effortlessly reason about sounds (e.g., pitch, loudness, or animal-sound associations) even without hearing them. In contrast, language models often lack such capabilities, limiting their effectiveness in multimodal interaction.
This benchmark provides a systematic way to measure whether LLMs… See the full description on the dataset page:
https://huggingface.co/datasets/HJOK/AuditoryBenchpp.