It's easy to assume a runtime safety check works because it catches an obvious NaN. Whether it actually catches the subtler failure, an action that's drifted somewhere the policy was never calibrated for, is a question you can only answer by measuring it. This is the small labelled set that lets you measure it.
It pairs real teleoperated robot actions (the in-distribution, "normal" class) with realistic faults injected into held-out real actions (the… See the full description on the dataset page:
https://huggingface.co/datasets/LaelaZ/vla-action-anomaly-eval.