← The frontier
Technology & AIAug 28, 2026

Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration

When training Mixture-of-Experts (MoE) language models with expert parallelism, all-to-all token dispatch and combine collectives can consume a substantial fraction of end-to-end training time.

When training Mixture-of-Experts (MoE) language models with expert parallelism, all-to-all token dispatch and combine collectives can consume a substantial fraction of end-to-end training time. In this work, we study communication-efficient MoE models (CE-MoE), in which we adopt…

The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.