LLM-as-judge safety checker: gpt-4o-mini judging flattened candidate action, user intent, policy text, proposed arguments, and evidence summaries.
AANA schema gate: consumes noisy aana.agent_tool_precheck.v1 events and returns accept, ask, defer, or refuse.
Source traces come from zake7749/Qwen-3.6-plus-agent-tool-calling-trajectory. Rows are… See the full description on the dataset page:
https://huggingface.co/datasets/mindbomber/aana-head-to-head-llm-judge-vs-aana.