NeurIPS 2026 Evaluations and Datasets Track submission (double-blind).
PLM-HalluBench is a benchmark for evaluating hallucination in protein language models (PLMs) — outputs that look like proteins but violate basic biophysics. It is organised around a three-level taxonomy that separates sequence-, structure-, and function-level failure modes, and it pairs a Factual track (BPHS against a… See the full description on the dataset page:
https://huggingface.co/datasets/plm-hallubench/plm-hallubench.