WildClean is a tabular cleaning benchmark with damage and silent-edit accounting — to our knowledge the first cleaning benchmark whose protocol charges systems for the clean cells they corrupt and for unattributed edits, not only for the errors they fix (a verified gap in the existing error-detection/repair benchmark literature, where recall-style scoring rewards over-correction).
It packages three complementary suites plus the entity vocabularies used by the reference… See the full description on the dataset page:
https://huggingface.co/datasets/ricalanis/wildclean.