Official v1 evaluation scenarios for DriftBench, a fully deterministic benchmark
for whether a system faithfully tracks belief drift, internal conflict, and identity
transition across a multi-turn conversation — scored with no LLM judge.
Code, validator, and full spec:
https://github.com/simon9679/driftbench (Apache-2.0)
Reliability research behind the benchmark:
https://github.com/simon9679/tbg-postmortem — a
negative-result study on a belief-memory… See the full description on the dataset page:
https://huggingface.co/datasets/simon9679/driftbench-v1.