TeleVRSLUBench is a spoken language understanding (SLU) benchmark that incorporates visual scene information and explicit reasoning processes for joint intent detection and slot filling.
The dataset is proposed in the paper:
Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding
It is, to the best of our knowledge, the first SLU benchmark that integrates scene-level visual context and… See the full description on the dataset page: https://huggingface.co/datasets/Tele-AI/TeleVRSLU.