FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
Karmesh Yadav*,
Yusuf Ali*,
Gunshi Gupta,
Yarin Gal,
Zsolt Kira
Current vision-language models (VLMs) struggle with long-term memory in embodied tasks. To address this, we introduce FindingDory, a benchmark in Habitat that evaluates memory-based reasoning across 60 long-horizon tasks.
In this repo, we release the FindingDory Subsampled Video Dataset. Each video contains 96 images… See the full description on the dataset page:
https://huggingface.co/datasets/yali30/findingdory-subsampled-96.