Compilation of ANnotated, Negation-Oriented Text-pairs
Dataset Card for CANNOT
Dataset Summary
CANNOT is a dataset that focuses on negated textual pairs. It currently
contains 77,376 samples, of which roughly of them are negated pairs of
sentences, and the other half are not (they are paraphrased versions of each
other).
The most frequent negation that appears in the dataset is verbal negation (e.g.,
will → won't), although it also contains pairs with… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/cannot-dataset.