Project Page | Paper | GitHub
LongMemEval-V2 (LME-V2) is an evaluation benchmark for long-term memory in web and enterprise agents. It contains 451 manually curated questions and 1,870 task trajectories drawn from customized WebArena-style and ServiceNow-style environments.
The benchmark evaluates five core memory abilities:
Static state recall: remembers important landmarks and page layouts.
Dynamic state tracking: understands how states change over time.… See the full description on the dataset page:
https://huggingface.co/datasets/xiaowu0162/longmemeval-v2.