HinDialectClassification
An MTEB dataset
Massive Text Embedding Benchmark
HinDialect: 26 Hindi-related languages and dialects of the Indic Continuum in North India
Task category
t2c
Domains
Social, Spoken, Written
Referencehttps://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-4839
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb