Existing LLM-based tools and coding agents respond to every issue and generate a patch for every case, even when the input is vague or their own output is incorrect. There are no mechanisms in place to abstain when confidence is low. BouncerBench checks if AI agents know when not to act.
This is one of 3 datasets released as part of the paper Is Your Automated Software Engineer Trustworthy?.
input_bouncerTasks on bug‐report text. The model decides if a report is… See the full description on the dataset page:
https://huggingface.co/datasets/uw-swag/input-bouncer.