Is Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older optimizers like gradient descent — and this matters for generalization.
Led by Sara Dragutinović and advised by Rajesh Ranganath
arxiv.org/abs/2603.00742