A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair.
When an LLM-based agent fails a task, a repair loop invokes an evaluator channel (unit test, linter, human rubric, etc.) to produce a diagnosis, which then guides the next repair attempt. AuditRepairBench reveals that the choice of evaluator channel is not merely a matter of cost or accuracy: different channels produce qualitatively different diagnoses for… See the full description on the dataset page:
https://huggingface.co/datasets/YueLinHu/AuditRepairBench.