Markdownified Ar5iv transforms academic papers from arXiv into clean, structured Markdown format consisting of 22.34B tokens across two splits. This dataset preserves th content while making it accessible for language model training on academic text.
License
C-UDA-1.0 (mirrors upstream Ar5iv licenses)… See the full description on the dataset page:
https://huggingface.co/datasets/marin-community/ar5iv-warning-markdown.