Benchmark for reaction plausibility, single-step and multistep retrosynthesis from Zagribelnyy et al. (2026).
It bundles two task families: reaction-level plausibility judgment and retrosynthesis target sets.
URSA-reaction-plausibility-bench-2026.csv — 1,000 reactions predicted by different models hand-labeled by expert oragnic chemists (500 plausible / 500 implausible), for benchmarking reaction-level… See the full description on the dataset page:
https://huggingface.co/datasets/insilicomedicine/URSA-benchmarking-sets.