This dataset accompanies our blog post Agentic coding improves ARC AGI 2 performance across models.
This contains the complete outputs from the relevant experiment runs; including the full prompts, llm responses and tool calls.
Description of the dataset
We provide data from experiment run using three models, in the folder interleaved_thinking_vs_plain_cot:
gpt_5_2_xhigh
gpt_oss_120b_high
minimax_m2_1
For each model, unless otherwise noted, we provide data for