"ArXiv-10" dataset consists of titles and abstracts extracted from 100k scientific papers on ArXiv, covering ten distinct research categories.
These categories includes subfields of computer science, physics, and mathematics.
To ensure consistency and manageability, the dataset consists of precisely 10k samples per category.
This dataset provides a practical resource for researchers and practitioners interested in LLM.
What is different about this dataset is the high… See the full description on the dataset page:
https://huggingface.co/datasets/effectiveML/ArXiv-10.