MT-AgentRisk is the first benchmark for evaluating multi-turn safety risks in tool-using LLM agents. As agents become increasingly capable, their safety lags behind — creating a widening capability-safety gap. This benchmark systematically scales up evaluations to multi-turn, tool realistic settings.
First, request access to the MT-AgentRisk dataset on… See the full description on the dataset page:
https://huggingface.co/datasets/CHATS-Lab/MT-AgentRisk.