CHOCOLATE is a benchmark for detecting and correcting factual inconsistency in generated chart captions. It consists of captions produced by six most advanced models, which are categorized into three subsets:
LVLM: GPT-4V, Bard (before Gemini)
LLM-based Pipeline: DePlot + GPT-4
Fine-tuned Model: ChartT5, MatCha, UniChart
The charts are from two datasets: VisText and the… See the full description on the dataset page:
https://huggingface.co/datasets/khhuang/CHOCOLATE.