This dataset contains embeddings from the
June 2024 Wikipedia dump
for the 11 most popular languages.
The embeddings are generated with the multilingual
BGE-M3 model.
The dataset consists of Wikipedia articles split into paragraphs,
and embedded with the aforementioned model.
To enhance search quality, the paragraphs are prefixed with their
respective article titles before embedding.
Additionally, paragraphs containing fewer than 100 characters… See the full description on the dataset page:
https://huggingface.co/datasets/Upstash/wikipedia-2024-06-bge-m3.