DocAtlas is a large-scale, high-fidelity multilingual OCR dataset covering 82 languages and 10 writing systems, built through model-free differential rendering. It provides precise structural annotations in a unified DocTag format encoding layout, text, and component types — without relying on any learned models for core annotation.This dataset powers the training of DocAtlas-DeepSeek, which achieves… See the full description on the dataset page:
https://huggingface.co/datasets/ahmedheakl/docatlas_instruct.