LoTTE
An MTEB dataset
Massive Text Embedding Benchmark
LoTTE (Long-Tail Topic-stratified Evaluation for IR) is designed to evaluate retrieval models on underrepresented, long-tail topics. Unlike MSMARCO or BEIR, LoTTE features domain-specific queries and passages from StackExchange (covering writing, recreation, science, technology, and lifestyle), providing a challenging out-of-domain generalization benchmark.
Task category
t2t
Domains
Academic, Web, Social
Reference… See the full description on the dataset page:
https://huggingface.co/datasets/mteb/LoTTE.