A benchmark dataset for evaluating how effectively AI agents can, through limited interaction, (i) infer users' design preferences and (ii) translate vague user intent into a specification sheet. Contains 50 software project proposals, 48 diverse simulated personas, 80 design questions per project, and benchmark results from 4 agents across 2 tasks.
Projects
50
Software proposals across 9 categories… See the full description on the dataset page:
https://huggingface.co/datasets/haowang94/specbench.