Can LLMs reliably extract structured medical timelines from unstructured records?
This dataset provides the golden ground truth, synthetic source documents, and pre-generated model outputs for benchmarking LLMs on medical chronology extraction — a critical task in medical-legal case review.
📦 GitHub (full code + evaluation pipeline): superinsight/superinsight-ai-benchmark
Tier
Models
Composite
F1
Hallucination… See the full description on the dataset page:
https://huggingface.co/datasets/Superinsight/medical-chronology-benchmark.