This release contains audited run-level results from a controlled study of execution-time safeguards for web agents under deceptive consumer interfaces.
The benchmark independently scores nominal completion (C) and trajectory safety (S): trustworthy completion (C=1,S=1), unsafe completion (C=1,S=0), safe non-completion (C=0,S=1), and unsafe failure (C=0,S=0).
One frozen vision-capable web-agent configuration
12… See the full description on the dataset page:
https://huggingface.co/datasets/deceptive-web-benchmark/execution-time-warnings-web-agents.