ProvTales is the first large-scale benchmark dataset for converting visual exploration histories (provenance graphs) into coherent data narratives.
It contains 22,560 provenance graph–data narrative pairs (11,280 constrained + 11,280 unconstrained), constructed from 365 real-world tabular datasets spanning 16 domains.
The dataset follows a narrative-first, graph-second construction pipeline:
target… See the full description on the dataset page:
https://huggingface.co/datasets/ZtZheng/ProvTales.