This page provides additional qualitative examples complementing the paper. It is created in response to reviewer feedback requesting more generated examples illustrating both successful reasoning and common failure modes across model families.
These examples visualize the RIG-Bench setting: a model receives visual context and a short instruction, infers the latent rule or target outcome, and synthesizes the answer directly as… See the full description on the dataset page:
https://huggingface.co/datasets/anonymous-submission-RIG-bench/RIG-Bench-Qualitative.