Judge-level evaluation data from a controlled comparison of automated
deep-research architectures over a 90-query manifest. These are the tidy frames
the paper's statistics are computed from: the reported numbers are recomputable
from the tables here using the analysis scripts in the code repository.
The paper compares eleven architectures on a single shared tool layer: a
single-pass baseline, eight orchestration… See the full description on the dataset page:
https://huggingface.co/datasets/PeterStrain77/bounded-returns-deep-research.