All evaluation outputs, reorganized by benchmark × run mode × model.
results/
├── ehr_bench/ # text-only EHR-Bench (1800 rows, 45 tasks)
│ ├── agentic/ # multi-turn tool-calling via deploy_agent.py
│ │ ├── smoke20/
/
│ │ ├── full1800// ← rollout outputs (results.jsonl, run.log, per-region shards)
│ │ └── scored/ ← evaluate_results.py output (WIP slot)
│ └── oneshot/… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/ClinSeek-Evaluation-Results.