日本語 | English
OmniDocBench-JASyn is a dataset for evaluating Japanese Document Parsing performance. It inherits the format of OmniDocBench and evaluates the end-to-end performance of VLMs on Japanese documents, covering OCR, layout analysis, table analysis, formula OCR, and reading order recognition.
Fully Synthetic Data : To ensure efficient and high-quality data, images and ground truth are generated by LLM-based code generation, followed by… See the full description on the dataset page:
https://huggingface.co/datasets/stockmark/OmniDocBench-JASyn.