This repo holds different preprocessed conditions of the same underlying
paper set as separate configs — load a specific one with
load_dataset("<repo_id>", "<config_name>").
152,520 rows, one row per paper — no concatenation of citing/cited pairs.
These are all of the unique papers that show up in the megapapers
project's (citing, cited) pairs: a Postgres DB scores every citation edge
between two Semantic Scholar papers with an… See the full description on the dataset page:
https://huggingface.co/datasets/hannahglz25/megapapers-pretrain.