SpokenCOCOI2ARetrieval
An MTEB dataset
Massive Text Embedding Benchmark
SpokenCOCO pairs MS COCO images with recordings of human speakers reading the corresponding English captions. This task uses the 5,000-image Karpathy test split with 25,031 spoken captions. Queries are images and the corpus contains spoken captions; the goal is to retrieve the recordings that describe each image.