TR-DocVQA-Synth is a large-scale synthetic Turkish Document Visual Question Answering dataset designed for training and evaluating multimodal models on Turkish business documents. The dataset contains 15,000 document images and 235,000 question-answer pairs generated from structured ground-truth records.
The dataset focuses on realistic Turkish document… See the full description on the dataset page: https://huggingface.co/datasets/Ethosoft/TR-DocVQA-Synth.