GOVA Word Acquisition — 3-Case Split by Language-Bias Regime
OctoBERT's grounded word-acquisition successes on GOVA (Flickr30k), partitioned
into three regimes by how much the answer depends on language vs. vision.
The two clean poles (Case 1 / Case 3) are defined by the intersection of two
models so they are model-agnostic, not an OctoBERT artifact: