1,971 Spanish premise–hypothesis pairs carrying two labels: the one derived automatically from
the discourse connector, and the one a majority of human annotators assigned after the connector was
removed. They agree on 972 pairs and disagree on 999.
An Analysis of the Performance of Large Language Models in Spanish NLI Datasets with Causal Relationships
Nicolás Pérez, Johan R. Portela, Ruben Manrique — Universidad de los Andes… See the full description on the dataset page:
https://huggingface.co/datasets/Flaglab/esnlir-human-majority.