Paper | Code
This repository contains the test split of the JSSODa dataset.
JSSODa (Japanese Simple Synthetic OCR Dataset) is constructed by rendering Japanese text generated by an LLM into images.
The images contain text written both vertically and horizontally, which is organized into one to four columns.
This dataset was introduced in our paper: "Evaluating Multimodal Large Language Models on Vertically Written Japanese Text".
The code… See the full description on the dataset page:
https://huggingface.co/datasets/llm-jp/JSSODa-test.