This is a pre-year-1900 dataset.
It is a mishmash of original vintage datasets, including: croqaz/vintage-v1, croqaz/vintage-v2, haykgrigorian/english-historical-corpus-1800-1875 and jbduran/think-dataset-clean. Only the highest quality texts are kept. Also, all entries that contain any of the ~800 banned words from the banned.txt file are dropped.
The texts are de-duplicated by lowercase + ignoring all white-spaces, but if one non-white-space character is… See the full description on the dataset page:
https://huggingface.co/datasets/croqaz/Sprocket-n-Say.