The second version of the BART pretraining corpus, focused on stripping low-quality text —
boilerplate and OCR corruption — out of
v1.
Lineage — three cumulative filtering stages over the same corpus:
Institutional Books 1.0 →
v1 →
v2 →
v3… See the full description on the dataset page:
https://huggingface.co/datasets/jbduran/bartholomew-dataset-v2.