Long-Horizon Personalization with Evolving Preferences
HorizonBench evaluates whether language models can track user preferences as they evolve across months of interaction. Each benchmark item is a 5-option multiple-choice question embedded within a conversation history averaging ~163K tokens. Pre-evolution preference values serve as hard-negative distractors, enabling diagnosis of belief-update failure: models retrieve the user's originally stated preference but fail… See the full description on the dataset page:
https://huggingface.co/datasets/stellalisy/HorizonBench.