Conversational experiences (context to distill into a compact KV representation) paired with
exact-match 4-option MCQ validation questions, derived from the distractor sessions of
LongMemEval (longmemeval_m, cleaned release). No LLM judge
needed: answers are single tokens, graded by letter match.
pooled3_experiences
88
conversations (~245k tok total), 3 questions… See the full description on the dataset page:
https://huggingface.co/datasets/hhy13/longmemeval-pooled-mcq.