This dataset is a query - context relevance correlation optimized for document retrieval in EN, CA and SP
Created from projecte_aina/RAG_Multilingual (train and validation splits, no test) and PaDaS-Lab/webfaq-retrieval datasets
It was created to fine tune embeddings models for use in Retrieval Augmented generation applications for these 3 languages
Context was limited to the previous and following sentences for the rellevant one… See the full description on the dataset page:
https://huggingface.co/datasets/langtech-innovation/trilingual_query_relevance.