HippoCamp is a benchmark for evaluating contextual agents in realistic, device-resident personal computing environments. Unlike agent benchmarks centered on web interaction, tool use, or generic software automation, HippoCamp focuses on multimodal file management over large personal file systems: agents must⦠See the full description on the dataset page: https://huggingface.co/datasets/MMMem-org/HippoCamp.