The STS-Hard dataset is a paraphrase detection test set derived from the STSBenchmark dataset. It was introduced as part of the PARAPHRASUS: A Comprehensive Benchmark for Evaluating Paraphrase Detection Models. The test set includes the paraphrase label as well as individual annotation labels from two annotators:
P1: The semanticist.
P2: A student annotator.
For more details, refer to the original paper that was presented at COLING 2025.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/sts-h-paraphrase-detection.