Evaluator-only, access-gated response corpora. Each complete entry contains 26 nonempty chat histories, 26 result JSON files, a manifest, and the run receipt when one survived. Benchmark source trees, tests, rubrics, and grader/oracle material are excluded.
Metric contract: pass@1 is the first-turn result. multi-turn-with-error-feedback@2 is the cumulative result after up to two sequential attempts on the same task, where the second… See the full description on the dataset page:
https://huggingface.co/datasets/TokenBender/glm47-aider-fixed26-responses.