Two evaluation dimensions your benchmarks are missing.
Pilot (n=44, 3 Anthropic model tiers, blind-scored): 100% categorical shift rate, zero reversals in 13 complete pairs (p=1.47e-03 Holm-Bonferroni corrected). Honored phrasing moves responses from generic-correct to expert-engaged. Effect is model-tier-agnostic in the pilot.
Cross-provider extension (9 models, 4 companies, 480 responses): Data collected, blind… See the full description on the dataset page:
https://huggingface.co/datasets/Wayfinder6/honored-ask-blaine-test.