ClawTrojan is a long-horizon agent safety evaluation dataset for studying
multi-stage, stealthy attack trajectories against tool-using agents. Instead of
modeling a sample as a single malicious prompt, ClawTrojan models an attack as a
trajectory that unfolds across user requests, tool outputs, downloaded files,
memory state, workspace files, and agent capabilities.
The dataset is part of the ClawShield project and is designed for research on
prompt injection detection… See the full description on the dataset page:
https://huggingface.co/datasets/zstanjj/ClawTrojan.