WorkSurface-Bench evaluates whether enterprise agents can route questions across
document retrieval (RAG), structured tables, and dependency graphs, acquire the
right evidence, and produce correct answers.
1,151 atomic tasks derived from 100 Workspace-Bench-Lite source tasks
5 persona-scoped workspaces
488 cross-surface, 279 table-only, 213 RAG-only, and 171 graph-only tasks
Cross-surface composition: 314 RAG+Graph, 100… See the full description on the dataset page:
https://huggingface.co/datasets/lhpku20010120/WorkSurface-Bench.