Document image parsing is challenging due to its complexly intertwined elements such as text paragraphs, figures, formulas, and tables. Dolphin addresses these challenges through a two-stage approach:
Dolphin achieves promising performance across diverse page-level and element-level parsing tasks while ensuring superior efficiency through its lightweight architecture and parallel parsing mechanism.
Our demo will be released in these days. Please keep tuned! 🔥
This model is released under the MIT License.
1@inproceedings{dolphin2025,
2 title={Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting},
3 author={Feng, Hao and Wei, Shu and Fei, Xiang and Shi, Wei and Han, Yingdong and Liao, Lei and Lu, Jinghui and Wu, Binghong and Liu, Qi and Lin, Chunhui and Tang, Jingqun and Liu, Hao and Huang, Can},
4 year={2025},
5 booktitle={Proceedings of the 65rd Annual Meeting of the Association for Computational Linguistics (ACL)}
6}