A multilingual counterfactual MCQ dataset built from rendered 3D object scenes.
Each row contains a rendered 3D scene image, two captions (original vs counterfactual), and a multiple-choice question probing whether a VLM follows the image or the misleading text.