Part of the Strategy Distillation experiment exploring whether synthesized reasoning strategies from large models can boost smaller models on AIME math.
pass@1: 60% (12/20)
20 heldout AIME questions (unseen by trace model), with distilled strategy. Thinking enabled.