Free-form model rollouts (generations) joined with blind LLM-judge verdicts
from a set of activation-steering and LoRA experiments on Llama-3.1-8B model
organisms. Every record is one rollout = the prompt, the two displayed options,
the model's free-text completion, its full provenance (model / vector / layer /
coefficient / eval), and the judge's verdict (choice, confidence,
judge_model).
All organisms are LoRA adapters on meta-llama/Llama-3.1-8B (the… See the full description on the dataset page:
https://huggingface.co/datasets/brikdavies/msm-graded-rollouts.