This dataset provides multilingual and code retrieval data for fine-tuning text embedding models. It is composed of high quality data sources with mined documents annotated with bi-encoder scores.
For each query, the 2048 closest documents are mined with snowflake-arctic-embed-l-v2.0 for MIRACL and MLDR and with gte-modernbert-base for CodeEditSearchTrain, and annotated with their bi-encoder similarity score. No false-negative filtering or cross-encoder annotation is… See the full description on the dataset page:
https://huggingface.co/datasets/lightonai/embeddings-fine-tuning-multilingual-unfiltered.