Textbooks scraped from libretext.org from the paper Language Models as Science Tutors.
This dataset uses the chapters from princeton-nlp/TextbookChapters and concatenates them by book and subject area to create long-context pre-training data.
If you use this dataset please cite:
@inproceedings{
chevalier2024language,
title={Language Models as Science Tutors},
author={Alexis Chevalier and Jiayi Geng and Alexander Wettig and Howard Chen and Sebastian Mizera and Toni Annala and Max Aragon and… See the full description on the dataset page:
https://huggingface.co/datasets/princeton-nlp/TextbooksBySubject.