Full agent transcripts for every published run of the Pareto Rail level-generation benchmark: models one-shot a browser rail-shooter level from a written theme, and visitors rank the results blind on the public site. The benchmark's methodology, per-run provenance manifests (rendered prompt, sealed commits, gate results, measured cost), and the level source produced by each of these transcripts live in the main repository:… See the full description on the dataset page:
https://huggingface.co/datasets/paulbatum/pareto-rail-rollouts.