If you are training base language models and want the cleanest sources available, OCD was built just for you.
Dataset Details
Dataset Description
It is without question that the quality of a language model rests on the quality of its training data. OCD is a meticulously curated and cleaned corpus of text documents, ensuring the highest quality text from a variety of sources. Part of this process includes manually… See the full description on the dataset page: https://huggingface.co/datasets/Daniel-P-Gonzalez/OCD.