This artifact reruns the AANA head-to-head agent-action comparisons on a second independent public tool-call source:
NousResearch/hermes-function-calling-v1
The goal is external validity, not an official leaderboard score. The source dataset provides function-calling conversations and tool schemas, but not human-reviewed safety labels. Safety labels here are policy-derived by the included transform script: