Ember is a lightweight optimizer for embedding tables and LM-head matrices. It replaces
Adam's dense first- and second-moment state on those layers — O(2VD) — with row/column
factored second moments, O(V+D): kilobytes of optimizer state instead of gigabytes, and no
sharding of token-table optimizer state in distributed setups.
Across supervised finetuning, RL, and pretraining, Ember matches Adam's validation loss on
these layers while carrying ~1500× less optimizer state.