GRUConcepts
Gated recurrent unit
1 · In one line
A GRU is a recurrent unit with reset and update gates.
1 · What it is
The result at time t is the hidden state h_t. When the reset gate is close to zero, the candidate ignores the previous hidden state. The update gate then interpolates between that candidate and the previous hidden state.
Cho and colleagues proposed an RNN Encoder-Decoder. A later comparison found GRU comparable to LSTM.
A recurrent wrapper can return either the last output or the full output sequence.
A GRU adds reset and update gates to a recurrent hidden unit.
Long gapsRecurrent networks can encounter vanishing and exploding gradient problems.
Smaller alternativeThe LSTM design described by the GRU paper has a memory cell and four gating units.
Follow the previous state through one GRU cell.
- 1 · receiveRead the current input and previous hidden state.
- 2 · resetUse the reset gate to decide whether the previous hidden state is ignored.
- 3 · proposeWhen the reset gate is close to zero, build the candidate from the current input while ignoring the previous hidden state.
- 4 · mixUse the update gate to blend the previous hidden state with the candidate.
- 5 · carryCompute the hidden state as (1 - z_t) ⊙ n_t + z_t ⊙ h_(t-1).
The proposed GRU uses two gating units, compared with four in the LSTM described by its paper.
| Who | What they ask | What it works with |
|---|---|---|
| Translation system | “Which phrase information should remain while reading the next symbol?” | An ordered source sequence |
| Motion modeller | “Which earlier positions help predict the next pose?” | A sequence of body coordinates |
| Speech team | “Which acoustic history matters for this frame?” | Ordered audio features |
solves
- The update gate controls how much information from the previous state carries over.
- The GRU update has an additive component from one time step to the next.
- The GRU paper describes two gating units, compared with four in its LSTM description.
doesn't solve
- GRU is not universally better than LSTM; comparative results depend on the task and model setup.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperLearning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation, Cho et al. · read 27 Sept 2026
- docsGRU, PyTorch · read 27 Sept 2026
- paperEmpirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling, Chung et al. · read 27 Sept 2026
- docsBase RNN layer, Keras · read 27 Sept 2026
- paperOn the difficulty of training Recurrent Neural Networks, Pascanu, Mikolov and Bengio · read 27 Sept 2026