A curated dataset of academic paper metadata sourced from arXiv, spanning key Computer Science and Machine Learning categories. Designed for research in text classification, topic modelling, summarisation, and scientometric analysis.
Dataset Description
Overview
The published snapshot (collected 2026-07-12) contains metadata for 2,019 papers — 2,500 were fetched and 481 cross-listed duplicates removed — across five core… See the full description on the dataset page: https://huggingface.co/datasets/gr8monk3ys/academic-papers-dataset.