Per GOVA word-acquisition works item (27,155): the fill distribution text-only
(blind) vs with the image for Qwen2.5-VL-7B-Instruct. Correctness-agnostic — it captures how
vision reshapes the prediction (passive language prior + active vision).
Organized into 5 splits by the vision-role taxonomy (3 cases; Case 1 & 3 sub-split
strong/weak; threshold high>0.5, boost>0.05, re-derivable from the stored probs):