A 30-task public sample of Health Optimization Bench,
a rubric-graded benchmark measuring how well frontier language models handle current clinical
evidence in preventive and optimization medicine. Three tasks from each of the benchmark's ten
micro benches.
The full benchmark is 977 authored tasks with 346 released across ten micro benches. On the
current leaderboard no model scores above 71 of 100 and the field spans 66 points. Rankings:… See the full description on the dataset page:
https://huggingface.co/datasets/Arcophos/health-optimization-bench-sample.