A coding-assistant memory benchmark for evaluating retrieval and recall over multi-session development histories. Think of it as LOCOMO for coding sessions.
The system ingests a series of Claude Code transcripts (JSONL) from realistic software projects and must answer questions that require remembering decisions, debugging steps, conventions, and temporal ordering across sessions.