Dataset Description
With LLMs, we can create a fully open-source Library of Alexandria.
As a first attempt, we have generated 650,000 unique textbook samples from a diverse span of courses, kindergarten through graduate school.
These are open source samples, which likely fall under the Llama-2 license. They were generated using the SciPhi repository.
All samples were created with TheBloke/Phind-CodeLlama-34B-v2-AWQ.
Lastly, I owe… See the full description on the dataset page:
https://huggingface.co/datasets/SciPhi/textbooks-are-all-you-need-lite.