Clustering of titles from arxiv. Clustering of 30 sets, either on the main or secondary category
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["ArXivHierarchicalClusteringS2S"])… See the full description on the dataset page:
https://huggingface.co/datasets/mteb/arxiv-clustering-s2s.