Dataset Card for Multilingual European Datasets for Sensitive Entity Detection in the Legal Domain
Dataset Summary
The dataset consists of 12 documents (9 for Spanish due to parsing errors) taken from EUR-Lex, a multilingual corpus of court
decisions and legal dispositions in the 24 official languages of the European Union. The documents have been annotated
for named entities following the guidelines of the MAPA project which foresees two
annotation level, a general and a… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/mapa.