Eval logs + exact oracle for zero-shot Symbolic Alchemy (no-summary / summary-replace / summary-augment).
Data: //{results.jsonl,summary.json,normalized.json,traj/}; oracle_all_episodes.json
(exact oracle, ALL 1000); difficulty_index.json; hard_set_20.json. Analysis: QWEN_FAILURE_MODES.md,
QWEN_SUMMARY_DRIFT.md. Full method/results below.
This repo now holds BOTH (1) the frontier-model eval baseline… See the full description on the dataset page:
https://huggingface.co/datasets/Racktic/alchemy-eval-logs.