This dataset is used for training the zilliz/semantic-highlight-bilingual-v1(
https://huggingface.co/zilliz/semantic-highlight-bilingual-v1) model for semantic highlighting in RAG (Retrieval-Augmented Generation) systems.
This dataset contains query-context pairs with relevance annotations for context spans. The annotations help identify which parts of a document are semantically relevant to a query… See the full description on the dataset page:
https://huggingface.co/datasets/zilliz/msmarco-context-relevance-with-think.