We sampled a set of ~5k charts and articles from Our World In Data to produce this evaluation set.
The set is comprised of 666 queries (text snippets from the articles in reference to the charts), and a total of 1k unique charts.
This particular dataset is a subsample of 1000 random charts from the full dataset which can be found here.
The text_description column contains OCR text extracted from the images using EasyOCR.… See the full description on the dataset page:
https://huggingface.co/datasets/jinaai/owid_charts_en.