This repository contains paraphrases for hypotheses in the ANLI training data.
Each original hypothesis has up to five paraphrases. Every retained paraphrase passed two automated checks. The first check tested semantic equivalence with the original hypothesis. The second check tested whether the relation between the premise and the paraphrase matched the original gold label.
The dataset contains 792,392 paraphrase rows from 161,395… See the full description on the dataset page: https://huggingface.co/datasets/Lidor-Mashiach/anli-paraphrase-bank.