PulmoBench: A Structured Clinical Reasoning Benchmark for Pulmonology
Overview
PulmoBench is a structured benchmark designed to evaluate Large Language Models (LLMs) in clinical risk stratification, safety-aware pulmonary reasoning, and zero-shot diagnostic generalization.
The benchmark features two evaluation tracks:
Core Benchmark (train, validation, test): Synthesized natural-language clinical vignettes with structured risk tiers, escalation flags, and… See the full description on the dataset page: https://huggingface.co/datasets/saibhossain/PulmoBench.