Benchmarking Realistic, Heterogeneous, and Evolving Long-Horizon Memory
RHELM is a benchmark for evaluating long-horizon memory capabilities in AI assistants.
Unlike benchmarks built around static dialogues, RHELM provides realistic,
heterogeneous, and temporally evolving memory sources, together with
challenging questions that require multi-hop reasoning, temporal synthesis, and
hallucination detection.