Measures whether a text-audio model can retrieve the call of a named animal. Each query is the taxon name of one held-out species or genus, and the model ranks 100 candidate taxa -- the queried taxon plus 99 distractors -- over an index of 1,024 field recordings, where a taxon is represented by every recording of that taxon. A taxon scores its best-matching recording and the 100 taxa are ranked by that score, so the… See the full description on the dataset page:
https://huggingface.co/datasets/myang333/BioVITAT2ARetrieval.