One row per run per task family: the paired pass-rate difference with a 95% confidence interval, the two-sided exact McNemar test, the full discordance breakdown, and a semantic-cluster-adjusted delta and interval.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this dataset:… See the full description on the dataset page:
https://huggingface.co/datasets/pbhappliedsystems/quant_eval_paired_degradation_statistics.