Can a model tell a dangerous agentic tool call from a safe one, given the surrounding
context? 3,000 held-out items, each a proposed tool call judged against the user's original
request and the agent's recent actions.
This is the evaluation set for
ProCreations/auto-1b and
ProCreations/auto-0.4b, and is excluded from
both training corpora (auto-1b-data,
auto) by content hash.