UnarXive 2024 is a large-scale, structured dataset of 2.3 million full-text arXiv papers (1991–2024), processed for use in NLP and information retrieval tasks. Each paper is provided in a structured JSONL format.
2.28 million structured papers across physics, CS, mathematics, and other fields
Logical section structure (Introduction, Methods… See the full description on the dataset page:
https://huggingface.co/datasets/ines-besrour/unarxive_2024.