This huggingface page contains data for the paper: Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
Evaluating the abilities of large language models (LLMs) for tasks that require long-term memory and thus long-context reasoning, for example in conversational settings, is hampered by the existing benchmarks, which often lack narrative coherence, cover narrow… See the full description on the dataset page:
https://huggingface.co/datasets/Mohammadta/BEAM.