10,000 pieces of news text from fancyzhx/ag_news with synthetically generated OCR mistakes.
The purpose of this is to mimic corrupt text that has been transcribed with OCR from old newspapers, where there are often lot's of errors. See biglam/bnl_newspapers1841-1879 for example. By synthetically creating it, we have the true ground truth, meaning we can use this as a source of truth for finetuning.
The corrupted text was generated using OpenAI's… See the full description on the dataset page:
https://huggingface.co/datasets/pbevan11/synthetic-ocr-correction-gpt4o.