The model is a hippocampal-entorhinal inspired world model that learns reusable transition structures from observation-only videos. It separates content-rich episodic representations from compact abstract dynamics, and uses velocity-like latent transitions for prediction and structural generalization across objects and scenes.
Please see the GitHub repository for installation, checkpoint loading, training, and evaluation instructions.
1@inproceedings{zhang2026structure,
2 title = {Structure Abstraction and Generalization in a Hippocampal-Entorhinal Inspired World Model},
3 author = {Zhang, Tianqiu and Lyu, Muyang and Liu, Xiao and Wu, Si},
4 booktitle = {Forty-third International Conference on Machine Learning},
5 year = {2026},
6 url = {https://openreview.net/forum?id=AYXgo5FjYz}
7}