Companion release for Certifying Compressed Language Models: An Audit and a
Statistical Toolkit.
The paper's headline audit finding is that 0 of the 16 eligible published
near-lossless claims release task-matched per-item outputs — the outputs a
third party would need to run the paired comparison themselves, covering the
tasks the claim is actually about. Three release outputs for other tasks only;
thirteen release none. This release is that… See the full description on the dataset page:
https://huggingface.co/datasets/AmoghSingh123/flipeval-artifacts.