The first 12-direction audio-video-text retrieval benchmark.
3,782 held-out triples scored on a shared gallery across all 6 single-modal
and 6 dual-modal directions — with human-corrected captions.
Standard audio–text and video–text benchmarks evaluate only single-modal
directions and never pool both video and audio in the same gallery.… See the full description on the dataset page:
https://huggingface.co/datasets/YunzeLiu/OmniRetriever-Bench.