Visual Fast Mapping (VFM) refers to the human ability to rapidly form new visual concepts from minimal examples based on experience and knowledge, a keystone of inductive capacity extensively studied in cognitive science. In the realm of computer vision, early endeavors tried to replicate this capability through one-shot learning methods yet achieving limited generalization. Despite the recent advancements in Visual Language Models… See the full description on the dataset page: https://huggingface.co/datasets/huaiming/VisualFastMappingBenchmark.