Strengthened the cross-source requirement that multiple-choice tasks are designed
around (selecting all correct options should require integrating both the SQL
result and the retrieved documents): 264 of 760 tasks were revised, with task IDs,
databases, and gold SQL unchanged. Documents-only accuracy drops from 61.5% to
38.7% while full-evidence accuracy stays at 80.6% (3 frontier models, strict
exact set match).
Diversified the number… See the full description on the dataset page:
https://huggingface.co/datasets/FDAbench2026/FDAbench-Full.