Vision version of the OpenReview-ICLR and arXiv PaperLens datasets.
Each row is one unique paper. We release all extracted papers — not every paper here is used in our downstream training/eval sets. The papers that are used are denoted by the references field, which lists every internal (release_name, release_split) pair the paper belongs to (a single paper can belong to multiple). reconstruction.py reads this field to materialize the original sharegpt data.json… See the full description on the dataset page:
https://huggingface.co/datasets/skonan/PaperLens-Vision.