Public results from Plumloom’s research on reliability in single-turn AI chat evaluations.
This dataset contains 45 controlled single-turn chat cases used to test whether repeated evaluation produces more stable model choices than one-shot evaluation. It also includes aggregate results from 38 measurement-reliability evaluation runs.
Private prompts, rubrics, model responses, and production execution details are intentionally excluded.