Open-ended evaluation dataset for measuring LLM performance on realistic
clinical laboratory and in vitro diagnostic (IVD) operations tasks —
pre-analytical, analytical, and post-analytical decisions grounded in
manufacturer instructions-for-use (IFUs), FDA submissions, and internal SOPs.
Companion evaluation harness: github.com//clinical-lab-bench
questions.jsonl — one question per line: question text, reference
answer, and a… See the full description on the dataset page:
https://huggingface.co/datasets/JasonTsai33/cliinical-lab-bench.