A synthetic chart question-answering dataset for training and evaluating
Vision-Language Models (VLMs). Each example is a rendered chart image paired with
a natural-language question and an answer that is computed from the underlying
data, so every label is ground truth by construction — not a language model's
guess.
Generated by the pipeline at
Code-based-Synthetic-Multimodal-Data-Generation.
Reference build: ~12,600 QA pairs… See the full description on the dataset page:
https://huggingface.co/datasets/dantelok/multidomain-chart-qa.