Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
library_of_congress_filtered – Dataset by common-pile | AlphaNeural AI
You can deploy this model and start earning money today!
common-pile
/
library_of_congress_filtered
like
0
text-generation
en
100K<n<1M
json
text
datasets
dask
mlcroissant
2506.05209
us
Views
No views yet
Model card
Files and Versions
Community
API
Library of Congress Description
The Library of Congress (LoC) curates a collection of public domain books called "Selected Digitized Books". We have downloaded over 130,000 English-language books from this public domain collection as OCR plain text files using the LoC APIs.
Dataset Statistics
Documents UTF-8 GB
129,052 35.6
License Issues
While we aim to produce datasets with completely accurate licensing information, license… See the full description on the dataset page:
https://huggingface.co/datasets/common-pile/library_of_congress_filtered
.