Paper | Code
This repository contains the train and validation splits of the JSSODa dataset.
JSSODa (Japanese Simple Synthetic OCR Dataset) is constructed by rendering Japanese text generated by an LLM into images.
The images contain text written both vertically and horizontally, which is organized into one to four columns.
This dataset was introduced in our paper: "Evaluating Multimodal Large Language Models on Vertically Written Japanese… See the full description on the dataset page:
https://huggingface.co/datasets/llm-jp/JSSODa.