Full-test-set reproduction of inference on MEWL (MachinE Word Learning,
Jiang et al., ICML 2023) — 9 word-learning
tasks x 600 test episodes, evaluated with 4 modern models (zero-shot).
Each row = one episode: 6 context images (each labeled with a novel-word
utterance), the query image, 5 candidate answers, the ground-truth
answer, and each model's prediction + correctness.
Subsets = the 9 task categories. Splits =… See the full description on the dataset page:
https://huggingface.co/datasets/guangliangliu/mewl-repro.