This contains answer_comparisons.json, for recreating the results of the paper titled "The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors".
This file contains ratings of answer pairs compared against each other. We used these ratings to investigate if models' mistaken predictions for erroneous student images align with ground truth for error-free student images.
Please consult the datacard for… See the full description on the dataset page:
https://huggingface.co/datasets/lucy3/aftermath_answer_comparison.