BIRD-Critic is the first SQL debugging benchmark designed to answer a critical question:
Can large language models (LLMs) fix user issues in real-world database applications? Each task in BIRD-CRITIC has been verified by human experts on the following dimensions:
Reproduction of errors on BIRD env to prevent data leakage.
Carefully curate test case functions for each task specifically.
Soft EX: This metric can evaluate SELECT-ONLY tasks.
Soft EX + Parsing:… See the full description on the dataset page:
https://huggingface.co/datasets/xx900221/bird-critic-1.0-flash-exp.