A small but clean Vietnamese-language Visual Question Answering dataset over 20 canonical Vietnamese dishes. Built as the Phase-1 corpus for the FoodLensVN project (final-year Deep Learning report).
5,572 (image, question, answer) rows — 4,460 train / 632 val / 480 test.
299 unique source images + 897 augmented variants (3× per source) → 1,196 image refs per variant.
Two image variants: raw/ (original aspect ratio, max edge ≤… See the full description on the dataset page:
https://huggingface.co/datasets/Tamir39/foodlensvn.