NutritionQA is a novel benchmark for understanding photos of nutrition labels, with practical applications like aiding users with visual impairments.
NutritionQA contains 50 photos of nutrition labels, each photo is paired with a descriptive question and a reasoning question (requires multi-hop reasoning).
The figure below shows that open VLMs perform poorly on NutritionQA, even after training on millions of images. Our code-guided synthetic data generation system can… See the full description on the dataset page:
https://huggingface.co/datasets/yyupenn/NutritionQA.