A memorization-proof, objectively-graded benchmark for one silent failure in ML-for-biology: a dominant confounder masking a true relatedness unit.
Frontier models know the recipes for leakage-aware dataset handling. They are not reliable at one specific, high-stakes thing: recovering a relatedness unit that a dominant confounder hides. The difficulty is that the correct method is non-obvious, and the obvious similarity signal points the wrong way.… See the full description on the dataset page:
https://huggingface.co/datasets/silterra/bioconfoundbench.