🧮 GDPval LLM Scaffolding Experiment (GPT-4o + Claude Sonnet)
Overview
This dataset contains model completions for a controlled behavioral experiment conducted by the Data Innovation Lab, UC Berkeley Haas.It explores how assistant scaffolding — structured planning and self-review guidance generated by Claude 3.5 Sonnet — affects the performance of GPT-4o on professional tasks drawn from the GDPval “gold” subset (OpenAI 2024).