This repository hosts a snapshot of the arXiv dataset, originally published on Kaggle, and re-uploaded to Hugging Face Datasets for easier access, versioning, and seamless integration with modern ML workflows.
The goal is to make large-scale arXiv metadata and paper content readily usable for LLM training, retrieval-augmented generation (RAG), citation analysis, and research analytics.
Depending on⦠See the full description on the dataset page:
https://huggingface.co/datasets/anuj0456/arxiv-dataset.