DigikalamagClustering
An MTEB dataset
Massive Text Embedding Benchmark
A total of 8,515 articles scraped from Digikala Online Magazine. This dataset includes seven different classes.
Task category
t2c
Domains
Web
Referencehttps://hooshvare.github.io/docs/datasets/tc
Source datasets:
PNLPhub/DigiMag
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb