A fixed, uniformly-drawn set of 4,000 real benign tool calls from Claude Code
sessions, built to measure false positives in agent action monitors. Each row is one
action a monitor would be asked to judge, plus a pointer into the source corpus that
reconstruct.py turns back into the transcript the monitor reads. (Two documented departures
from live shape — the compaction cut and mid-turn truncation of batched API… See the full description on the dataset page:
https://huggingface.co/datasets/Syghmon/swe-benign-4000.