Search 3.2M models and datasets…
⌘K
Chat
Models
Datasets
Deploy
Pricing
Docs
Chat
Models
Datasets
Deploy
More
sts22-crosslingual-sts – Dataset by mteb | AlphaNeural AI
Is this your dataset? Claim it with the Hugging Face account that owns it.
mteb
/
sts22-crosslingual-sts
like
0
sentence-similarity
semantic-similarity-scoring
human-annotated
multilingual
mteb/sts22-crosslingual-sts
ara
cmn
deu
eng
fra
ita
pol
rus
spa
tur
unknown
10K<n<100K
json
text
datasets
Views
No views yet
Dataset card
Files and Versions
Community
Use
Use this dataset
STS22.v2 An MTEB dataset Massive Text Embedding Benchmark
SemEval 2022 Task 8: Multilingual News Article Similarity. Version 2 filters updated on STS22 by removing pairs where one of entries contain empty sentences.
Task category t2t
Domains News, Written
Reference
https://competitions.codalab.org/competitions/33835
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code: import mteb
task =… See the full description on the dataset page:
https://huggingface.co/datasets/mteb/sts22-crosslingual-sts
.