LITMUS is the first benchmark specifically designed to evaluate the behavioral safetyof LLM-based autonomous agents operating in real OS environments. It addresses two critical gaps left by prior work:
Semantic-only evaluation misses physical-layer harms. Existing benchmarks judge safety solely by an agent's verbal output, overlooking *Execution Hallucination… See the full description on the dataset page:
https://huggingface.co/datasets/AlienZhang1996/LITMUS.