25,756 graded outputs from nine pinned open-weight models reading real SEC DEF 14A compensation
tables, each one scored against the filer's own XBRL tag rather than a model judge, and each one
carrying the token-level uncertainty the model reported when it answered.
AI disclosure: the research is the author's; this text was drafted with AI assistance and reviewed by the author. The model, and the conflict… See the full description on the dataset page:
https://huggingface.co/datasets/NMAIResearch/model-dependency.