ProVision is a benchmark dataset for evaluating state-of-the-art multi-modal language models (MLLMs) across diverse tasks such as science, coding, creative writing, information extraction, perception, knowledge, arts, planning, and mathematics. The dataset aggregates chat instances with associated images, reference answers from gpt-4o, and meta-information (e.g. challenge difficulty, category labels, language, and interconnect decisions) to facilitate comprehensive… See the full description on the dataset page:
https://huggingface.co/datasets/HelloKKMe/ProBench.