Mixture of experts
A mixture-of-experts layer uses a learned gate to select a sparse combination of expert subnetworks for each input.
A trainable gating network determines a sparse combination of experts for each input. Each expert network computes an output. The layer sums selected expert outputs after multiplying each by its gating value.
Switch Transformer routes each token to one expert. GShard uses a gating function to route each token to the best two experts. Mixtral has eight feed-forward experts in each layer and selects two for each token.
Sparse routing can create load imbalance between experts. Expert capacity sets the number of tokens each expert computes. Experts placed on different devices also create communication cost between devices.
A model can add expert parameters without using every expert for every input.
Follow one token through a sparse expert layer.
- 1 · scoreA trainable gating network produces expert weights for the input.
- 2 · selectThe gate keeps a sparse combination of experts.
- 3 · computeThe selected expert subnetworks process the input.
- 4 · combineThe layer adds the selected expert outputs using the gate values.
Sparse routing means an input uses only part of the expert set.
| Who | What they ask | What it works with |
|---|---|---|
| Language-model team | “Which experts should process this token?” | Router scores for the token |
| Training engineer | “Are tokens distributed across experts?” | Expert load and capacity |
| Systems engineer | “How much expert communication occurs across devices?” | Routed token transfers |
- The original sparsely gated MoE layer used up to thousands of feed-forward subnetworks.
- Switch Transformer routes each token to one expert.
- GShard uses a gating function to route each token to the best two experts.
- Expert routing can create load imbalance.
- Expert capacity can cause tokens to overflow an expert's available buffer.
- Distributed experts require communication between devices.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperOutrageously Large Neural Networks, Shazeer et al. · read 28 Sept 2026
- paperSwitch Transformers, Fedus, Zoph and Shazeer · read 28 Sept 2026
- paperGShard, Lepikhin et al. · read 28 Sept 2026
- paperST-MoE, Zoph et al. · read 28 Sept 2026
- paperMixtral of Experts, Jiang et al. · read 28 Sept 2026