LPMusicCapsMTTA2TRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
LLM-generated pseudo captions for 10-second music clips from the MagnaTagATune dataset. Captions were produced by prompting a large language model with the human-annotated tags of each clip, giving four differently-styled captions per clip. Complements MusicCaps, whose captions are human-written and whose audio comes from AudioSet.