NLPJournalAbsArticleRetrieval.V2
An MTEB dataset
Massive Text Embedding Benchmark
This dataset was created from the Japanese NLP Journal LaTeX Corpus. The titles, abstracts and introductions of the academic papers were shuffled. The goal is to find the corresponding full article with the given abstract. This is the V2 dataset (last updated 2025-06-15).