VoxPopuli-NER is derived from the VoxPopuli corpus and specifically enhanced for
Named Entity Recognition (NER) tasks focusing on political and geographical entities.
It includes 879 audio samples, annotated with 2469 unique entity types. The dataset consists of the English part of the test set of VoxPopuli.
See full details in the WhisperNER paper.
If you find this usful, please cite the following works:
@article{ayache2024whisperner… See the full description on the dataset page:
https://huggingface.co/datasets/aiola/Voxpopuli_NER.