This dataset contains the raw output of the experiments of our paper The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?.
mcq/: Raw JSONL files of all runs with the LM Eval Harness Fork here.
mwe/: Model-Written Eval Suite, both multiple choice mcq and open-ended formats, obtained with the codebase of the… See the full description on the dataset page:
https://huggingface.co/datasets/hot-mess/hot-mess-data.