30,000 generations: 5 models (1.7B-3.8B) x 4 quant levels (FP16/Q8_0/Q4_K_M/
Q3_K_M) x 500 machine-checkable structured-output & tool-call tasks x 3 seeds,
run with llama.cpp on free-tier T4s, scored deterministically (no LLM judges),
aggregated with paired bootstrap 95% CIs.
Headline: Q8_0 showed zero significant regressions across 75 comparisons;
Q3_K_M significantly degrades schema compliance in 3/5 models and collapses… See the full description on the dataset page:
https://huggingface.co/datasets/kash-on-the-dash/quantone.