English | 简体中文
EgoMemo is a training-free streaming system for first-person (egocentric) video understanding and proactive assistance. As video streams in time-slice by time-slice, the system transcribes the scene into searchable long-term memory (captions + knowledge graph + vector store) while, in parallel, deciding on its own whether to proactively alert the user and answering user questions in real time. It runs in both CLI and… See the full description on the dataset page:
https://huggingface.co/datasets/SitongGong/egomemo_demo.