This dataset was built upon Syntec information retrieval dataset, negative samples were created using BM25.
Please refer to our paper for more details.
If you use this dataset in your work, please consider citing:
@misc{ciancone2024extending,
title={Extending the Massive Text Embedding Benchmark to French},
author={Mathieu Ciancone and Imene Kerboua and Marion Schaeffer and Wissam Siblini},
year={2024},
eprint={2405.20468}… See the full description on the dataset page:
https://huggingface.co/datasets/lyon-nlp/mteb-fr-reranking-syntec-s2p.