A trace-grounded benchmark for evaluating tool-using LLM agents over live
Model Context Protocol (MCP) servers. Tasks are generated forward: a
goal-seeded explorer agent drives live MCP tools into a recorded execution
trace, which is distilled into a TaskSpec scored on effects — path-agnostic
effect checkpoints (with equivalence sets, argument predicates), value_produced
checkpoints, minefields (must-not effects), and a partial order — never on
answer-matching.… See the full description on the dataset page:
https://huggingface.co/datasets/anonsubmitter/DynamicMCPBench.