Benchmark data for the paper "Securing LLM-Agent Long-Term Memory Against Poisoning: Non-Malleable, Origin-Bound Authority with Machine-Checked Guarantees." This dataset holds the scenarios and the
result logs. The full code, harness, and TLA+ formal model live in the companion
GitHub repository.
Code / harness / formal model (GitHub): [
https://github.com/yedidel/mem-inv-bench ]
LLM agents with persistent memory can be poisoned:… See the full description on the dataset page:
https://huggingface.co/datasets/anonymos-2321135/MEM-INV-Bench.