CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
Dataset Summary
CC-OCR V2 is a comprehensive and challenging OCR benchmark tailored to real-world document processing. It focuses on practical enterprise document processing tasks and incorporates hard and corner cases that are critical yet underrepresented in prior benchmarks.
The dataset comprises 7,093 high-difficulty samples covering 5 major OCR-centric tracks: Text… See the full description on the dataset page: https://huggingface.co/datasets/Eioss/CC-OCR-V2.