PressBooks is a searchable catalog of over 8,000 open access books.
To collect openly licensed content from PressBooks we construct a search query to retrieve URLs for all books written in English and listed as public domain or under CC BY or CC BY-SA licenses.
For each matched book, we collect its contents directly from the publicly available web version provided by PressBooks.
Per-document license information is available in the license entry of… See the full description on the dataset page:
https://huggingface.co/datasets/common-pile/pressbooks_filtered.