This dataset contains anchor papers with their top-K most similar (positive) and most dissimilar (negative) papers based on SPECTER2 embeddings.
anchor_id: Unique identifier for the anchor paper
anchor_title: Title of the anchor paper
anchor_abstract: Abstract of the anchor paper
positive_pool: List of 5 most similar papers, each as [id, title, abstract]
negative_pool: List of 5 most dissimilar… See the full description on the dataset page:
https://huggingface.co/datasets/Jerjes/neuro-specter2-sample-data.