CLIR-Bench is an evidence-auditable benchmark for multimodal question answering over irregular clinical time series, constructed from de-identified ICU records in MIMIC-IV. It contains 6,600 multiple-choice QA instances spanning 11 clinical variables and intervention signals, organized into four core capability dimensions: temporal understanding, reasoning, forecasting, and decision-making. Each question is explicitly linked to timestamp-level evidence and task-specific answer derivation rules… See the full description on the dataset page:
https://huggingface.co/datasets/zyadasz/CLIR-Bench.