The dataset was generated using LangChain's wrapper around GPT-4o-mini, with additional randomization performed by GPT-4.5.
Randomization has been introduced with token shuffling and python Faker library.
The goal was to create a dataset that is 90% clean, while intentionally introducing 10% of samples with OCR-like noise and artifacts. These imperfections are characterized by:
Excessive spacing between words (e.g., three or more spaces instead of one),
Erratic line breaks… See the full description on the dataset page:
https://huggingface.co/datasets/kris-szczepaniak/synthetic-documents-8k.