Clustering of titles from 199 subreddits. Clustering of 25 sets, each with 10-50 classes, and each class with 100 - 1000 sentences.
Domains
Web, Social, Written
Reference
https://arxiv.org/abs/2104.07081
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["RedditClustering.v2"])… See the full description on the dataset page:
https://huggingface.co/datasets/mteb/reddit-clustering.