Work in Progress (WIP)
This is an early publication. We are actively working on improving OCR quality and expanding coverage.
This dataset contains page-level OCR output (as Markdown) from a public release of documents related to Jeffrey Epstein / the Epstein case.
Each Markdown file represents one scanned page converted to text using an automated OCR pipeline. The dataset is designed for: