An evaluation framework that measures whether large language models augment or substitute for human cognition.
This dataset accompanies the NeurIPS 2026 Evaluations & Datasets Track submission "The Pro-Worker AI Benchmark: Measuring Whether Large Language Models Augment or Replace Human Intelligence".
prompts/layer1_behavioral/
200 single-turn behavioral probes across 10 dimensions (10 YAML… See the full description on the dataset page:
https://huggingface.co/datasets/pwb-anon-2026/pro-worker-ai-benchmark.