STSBenchmark.v2
An MTEB dataset
Massive Text Embedding Benchmark
Semantic Textual Similarity Benchmark (STSbenchmark) dataset. This version removes duplicate sentence pairs from the validation and test splits when the same order-insensitive pair appears more than once in an evaluation split or appears in another split.
Task category
STS (text-to-text)
Domains
Blog, News, Written
Reference
Machine translated multilingual STS benchmark dataset.
Source datasets:… See the full description on the dataset page:
https://huggingface.co/datasets/mteb/STSBenchmarkv2.