This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
This dataset contains pairs of prompts and classifications evaluating the safety of AI agent instructions in software development contexts. Each sample presents a scenario where an agent is asked to perform a task, labeled as either 'benign' for safe operations or 'suspicious' for actions involving security violations, data exfiltration, or privilege… See the full description on the dataset page:
https://huggingface.co/datasets/melanieyes/adaption-ai-agent-safety-prompts.