A high-quality multimodal dataset derived from the official paper “DeepSeek-OCR: Contexts Optical Compression for Ultra-Long Document Understanding”. This dataset contains 2,200 structured image–text pairs covering diverse document types—including academic papers, financial reports, textbooks, slides, newspapers, charts, chemical formulas, and geometric figures—extracted and reformatted to support end-to-end OCR, layout parsing, and vision-language pretraining… See the full description on the dataset page:
https://huggingface.co/datasets/amishor/DeepSeek-OCR.