Concepts

Recurrent neural networks

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

At each time step, a recurrent neural network updates its hidden state from the previous hidden state and the current input.

1 · What it is

An RNN receives an input tensor and an initial hidden state. At time t, it updates the hidden state from the current input and the previous hidden state. The next update receives that hidden state.

Reusing this update lets one model operate on sequences whose lengths vary. Recurrent networks have been used for sequential data such as text and speech.

RNN training can encounter vanishing and exploding gradients.

2 · Why it exists

The meaning of one item in a sequence often depends on what came before it.

Variable lengthSequences such as sentences and recordings do not all contain the same number of steps.
Order mattersRearranging the same items can change the sequence being modelled.
Shared ruleA sequence model needs a repeatable update rather than a different set of weights for every possible position.
3 · How it works

Follow a hidden state through three words.

Each sequence element is processed with the same documented state-update function.
  1. 1 · receiveRead the current input and the hidden state from the previous time step.
  2. 2 · updateCombine both with learned input-to-hidden and hidden-to-hidden weights, then apply a non-linearity.
  3. 3 · carryPass the new hidden state to the next time step.
  4. 4 · emitReturn every step's output or only the final state, depending on the task.

The recurrence is the state update repeated for each sequence element.

4 · Where it's used
WhoWhat they askWhat it works with
Speech team“Which sound is likely after the frames heard so far?”A sequence of audio features
Forecasting team“What value is likely to follow this history?”Ordered measurements over time
Language team“Which symbol is likely to come next?”Earlier symbols in the sequence
5 · What it solves, and what it doesn't
solves
  • An RNN can operate on a variable-length sequence.
  • The hidden state lets later outputs depend on earlier inputs.
  • A bidirectional variant can process a finite sequence in both directions.
doesn't solve
  • RNN training can encounter vanishing and exploding gradients.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsRNN, PyTorch · read 27 Sept 2026
  2. docsBase RNN layer, Keras · read 27 Sept 2026
  3. paperLearning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation, Cho et al. · read 27 Sept 2026
  4. paperOn the difficulty of training Recurrent Neural Networks, Pascanu, Mikolov and Bengio · read 27 Sept 2026
  5. paperDeep learning, LeCun, Bengio and Hinton · read 27 Sept 2026