SudoSet: A Contextual Authorization Benchmark for LLM Agents
SudoSet measures whether frontier LLMs honor classical authorization primitives — information-flow tags on observation provenance and capabilities on session scopes — when a runtime exposes them through the prompt context. It covers three attack families along a paired-contrast structure: confused deputy, source-label indirect prompt injection, and scope-check (transitive) IPI.