Welcome to the Open License Corpus (OLC), a 228B token corpus for training permissively-licensed language models.
Disclaimer: OLC should not be considered a universally safe-to-use dataset. We encourage users of OLC to consult a legal professional on the suitability of each data source for their application.
Legal
Case Law, Pile of Law (PD subset)
Public… See the full description on the dataset page:
https://huggingface.co/datasets/kernelmachine/open-license-corpus.