This dataset contains a probabilistic sample of publicly available PubMed metadata sourced from the National Library of Medicine (NLM).
If you're looking for the precomputed embedding vectors (MedCPT) used in our work Efficient and Reproducible Biomedical Question Answering using Retrieval Augmented Generation, they are available in a separate dataset: slinusc/PubMedAbstractsSubsetEmbedded.
Each entry in the dataset includes:… See the full description on the dataset page:
https://huggingface.co/datasets/slinusc/PubMedAbstractsSubset.