A benchmark dataset for evaluating long-term memory capabilities in conversational AI systems. It is part of EverMemBench, the first benchmark designed for long-horizon collaborative memory, introduced in the paper Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues — accepted at KDD 2026 (Oral).
Multi-turn group dialogues spanning ~250… See the full description on the dataset page:
https://huggingface.co/datasets/EverMind-AI/EverMemBench-Dynamic.