Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
s2orc_arxiv – Dataset by jedibear | AlphaNeural AI
You can deploy this model and start earning money today!
jedibear
/
s2orc_arxiv
like
0
text-generation
summarization
feature-extraction
en
1M<n<10M
text
us
s2orc
arxiv
scientific-papers
nlp
research
Views
No views yet
Model card
Files and Versions
Community
API
S2ORC ArXiv
A subset of the Semantic Scholar Open Research Corpus (S2ORC) filtered to ArXiv papers. Contains 2.58 million parsed scientific papers with full text, abstracts, structured sections, figures, and citation metadata.
Dataset Summary
Statistic Value
Total papers 2,579,762
Total size ~266 GB
Format Parquet
Split train
Dataset Structure Content Fields
Field Type Description
title string Paper title
abstract… See the full description on the dataset page:
https://huggingface.co/datasets/jedibear/s2orc_arxiv
.