A sampled subset (~3GB) of the Clotho v2.1 development split, packaged for quick experimentation with audio-text retrieval pipelines.
This dataset is a convenience subset of the Clotho audio captioning dataset, created for rapid prototyping and testing of audio-text retrieval models (e.g., CLAP fine-tuning) on limited compute.
Source: Clotho v2.1 (development split)
Original Authors: K. Drossos, S. Lipping, T.… See the full description on the dataset page:
https://huggingface.co/datasets/LakoreAI/clotho-dev-sample.