Technology & AIAug 7, 2026
Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers
Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head.
Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on $(a+b) \bmod 113$ grok and later lose generalization. Across five seeds the selected AdamW…
Sign in to learn & save →
The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.