Concepts

Mixture of experts

3 min readadvancedUpdated 28 Sept 2026
1 · In one line

A mixture-of-experts layer uses a learned gate to select a sparse combination of expert subnetworks for each input.

1 · What it is

A trainable gating network determines a sparse combination of experts for each input. Each expert network computes an output. The layer sums selected expert outputs after multiplying each by its gating value.

Switch Transformer routes each token to one expert. GShard uses a gating function to route each token to the best two experts. Mixtral has eight feed-forward experts in each layer and selects two for each token.

Sparse routing can create load imbalance between experts. Expert capacity sets the number of tokens each expert computes. Experts placed on different devices also create communication cost between devices.

2 · Why it exists

A model can add expert parameters without using every expert for every input.

Sparse useThe gate selects a sparse combination of experts for each example.
Uneven trafficA load-balancing loss encourages experts to receive roughly equal numbers of training examples.
Device limitsExpert capacity limits how many tokens each expert processes in one batch.
3 · How it works

Follow one token through a sparse expert layer.

Sparse mixture-of-experts routing A token is scored by a router. The top two routes select experts one and four. Their weighted results combine into one output while experts two and three remain inactive.
The learned gate selects a sparse set of experts, and their weighted outputs are combined.
  1. 1 · scoreA trainable gating network produces expert weights for the input.
  2. 2 · selectThe gate keeps a sparse combination of experts.
  3. 3 · computeThe selected expert subnetworks process the input.
  4. 4 · combineThe layer adds the selected expert outputs using the gate values.

Sparse routing means an input uses only part of the expert set.

4 · Where it's used
WhoWhat they askWhat it works with
Language-model team“Which experts should process this token?”Router scores for the token
Training engineer“Are tokens distributed across experts?”Expert load and capacity
Systems engineer“How much expert communication occurs across devices?”Routed token transfers
5 · What it solves, and what it doesn't
solves
  • The original sparsely gated MoE layer used up to thousands of feed-forward subnetworks.
  • Switch Transformer routes each token to one expert.
  • GShard uses a gating function to route each token to the best two experts.
doesn't solve
  • Expert routing can create load imbalance.
  • Expert capacity can cause tokens to overflow an expert's available buffer.
  • Distributed experts require communication between devices.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperOutrageously Large Neural Networks, Shazeer et al. · read 28 Sept 2026
  2. paperSwitch Transformers, Fedus, Zoph and Shazeer · read 28 Sept 2026
  3. paperGShard, Lepikhin et al. · read 28 Sept 2026
  4. paperST-MoE, Zoph et al. · read 28 Sept 2026
  5. paperMixtral of Experts, Jiang et al. · read 28 Sept 2026