Search 3.2M models and datasets…
⌘K
Chat
Models
Datasets
Deploy
Pricing
Docs
Chat
Models
Datasets
Deploy
More
stsbenchmark-sts – Dataset by mteb | AlphaNeural AI
Is this your dataset? Claim it with the Hugging Face account that owns it.
mteb
/
stsbenchmark-sts
like
0
sentence-similarity
semantic-similarity-scoring
human-annotated
translated
eng
unknown
1K<n<10K
json
text
datasets
pandas
polars
mlcroissant
2502.13595
2210.07316
us
mteb
text
Views
No views yet
Dataset card
Files and Versions
Community
Use
Use this dataset
STSBenchmark An MTEB dataset Massive Text Embedding Benchmark
Semantic Textual Similarity Benchmark (STSbenchmark) dataset.
Task category t2t
Domains Blog, News, Written
Reference
https://github.com/PhilipMay/stsb-multi-mt/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code: import mteb
task = mteb.get_tasks(["STSBenchmark"]) evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page:
https://huggingface.co/datasets/mteb/stsbenchmark-sts
.