This is a subset of the Multilingual Spoken Word Corpus dataset, which is built specifically for the Few-shot Class-incremental Learning (FSCIL) task.
A total of 15 languages are chosen, split into 5 base languages (English, German, Catalan, French, Kinyarwanda) and 10 incrementally learned languages (Persian, Spanish, Russian, Welsh, Italian, Basque, Polish, Esparanto, Portuguese, Dutch).
The FSCIL task entails first training a model using abundant training data on words from the 5 base… See the full description on the dataset page:
https://huggingface.co/datasets/NeuroBench/mswc_fscil_subset.