A deterministic set of 20 sanitized synthetic agent traces designed to exercise the Open Agent Failure Atlas detectors across safe, scope, approval, injection, recovery, efficiency, secret, and traversal cases.
This dataset is a software smoke test. A perfect score on it is not evidence of real-world precision, recall, robustness, or model safety. The examples are deliberately constructed around… See the full description on the dataset page:
https://huggingface.co/datasets/solsticestudioai/agent-failure-atlas-benchmark.