This dataset is basically mapping of anchor-positive pair chunks.
It consists of consolidated "title"+"summary"+"inclusion_criteria" for each CT idi.e. clinical trials(unique "nctId") as chunks.
Each of the aformentioned chunks have 4 questions(anchors) branched to it(here 1-to-1 normalized mapping of those).
The data is particularly useful in order to fine tune an embedding model for CT domain. This further helps in (i)ranked retreival genesis of CT RAGs… See the full description on the dataset page:
https://huggingface.co/datasets/vab46/Clinical_trials_anchor-positive-pairs_EmbeddingModel-data.