This project contains the behavior dataset in BrowserART, a red teaming test suit tailored particularly for browser agents.
For safety reasons, large language models (LLMs) are trained to refuse harmful user instructions, such as assisting dangerous activities. We study an open question in this work: Can the desired safety refusal, typically enforced in chat… See the full description on the dataset page:
https://huggingface.co/datasets/ScaleAI/BrowserART.