Per-question per-mode generation traces (model outputs + uncertainty
features + LLM-judge correctness labels) produced by the QEMR paper's
multi-VLM × multi-evidence-mode sweep.
These traces are what feeds python3 reproduce.py --analysis in the
QEMR code repository. Together
the two reproduce all paper tables on CPU in minutes — without
re-running GPU inference, without re-rendering HM3D scenes, without
any Habitat-Sim install.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/QEMR-2026/qemr-traces.