This dataset has been built from the Syntec Collective bargaining agreement. Its purpose is information retrieval.
The dataset is rather small. It is intended to be used only as a test set, for fast evaluation of models.
It is split into 2 subsets :
queries : it features 100 manually created questions. Each question is mapped to the article that contains the answer.
documents : corresponds to the 90 articles from… See the full description on the dataset page:
https://huggingface.co/datasets/lyon-nlp/mteb-fr-retrieval-syntec-s2p.