The 70-task dataset behind MCP-Bench —
a reliability benchmark for open-source LLM agents using Anthropic's
Model Context Protocol (MCP).
Tasks: 70
MCP servers covered:
fetch: 13
filesystem: 14
github: 10
memory: 10
postgres: 10
sqlite: 13
Difficulty:
medium: 34
easy: 24
hard: 12
Skill (RQ1/RQ2 stratifier):
selection: 40
composition: 23
recovery: 7
Each task is a JSON object with these… See the full description on the dataset page:
https://huggingface.co/datasets/keerthanaSubru57/mcp-bench.