When several large language models advise on a task that has no verifiable correct
answer (strategy, ethics, policy, crisis trade-offs), "which model is right" is the
wrong question. The useful question is how, and how much, the models differ in the
structure of their judgment. CM-RG measures exactly that. It adapts George Kelly's
Personal Construct Psychology (1955): each model writes a free-text advisory
response, elicits its own bipolar… See the full description on the dataset page:
https://huggingface.co/datasets/sergeydolgov/cross-model-repertory-grid.