Ablating Optimizer Ideas Around MuonAdamW for Transformer Pretraining
Overview
27 optimizer experiments on a 50M-parameter transformer, systematically testing whether anything can beat a well-tuned MuonAdamW baseline. We tried curvature-aware momentum, spectral boosting via power iteration, dual-timescale EMAs, variance-adaptive scaling, sign-coherence weighting, and more — drawing from recent papers (NAMO, AdEMAMix, Muon-VS, ROOT) and first-principles linear algebra.… See the full description on the dataset page: https://huggingface.co/datasets/mishig/autoresearch-optimizer-findings.