Concepts
Recurrent neural networks
1 · In one line
At each time step, a recurrent neural network updates its hidden state from the previous hidden state and the current input.
1 · What it is
An RNN receives an input tensor and an initial hidden state. At time t, it updates the hidden state from the current input and the previous hidden state. The next update receives that hidden state.
Reusing this update lets one model operate on sequences whose lengths vary. Recurrent networks have been used for sequential data such as text and speech.
RNN training can encounter vanishing and exploding gradients.
The meaning of one item in a sequence often depends on what came before it.
Variable lengthSequences such as sentences and recordings do not all contain the same number of steps.
Order mattersRearranging the same items can change the sequence being modelled.
Shared ruleA sequence model needs a repeatable update rather than a different set of weights for every possible position.
Follow a hidden state through three words.
- 1 · receiveRead the current input and the hidden state from the previous time step.
- 2 · updateCombine both with learned input-to-hidden and hidden-to-hidden weights, then apply a non-linearity.
- 3 · carryPass the new hidden state to the next time step.
- 4 · emitReturn every step's output or only the final state, depending on the task.
The recurrence is the state update repeated for each sequence element.
| Who | What they ask | What it works with |
|---|---|---|
| Speech team | “Which sound is likely after the frames heard so far?” | A sequence of audio features |
| Forecasting team | “What value is likely to follow this history?” | Ordered measurements over time |
| Language team | “Which symbol is likely to come next?” | Earlier symbols in the sequence |
solves
- An RNN can operate on a variable-length sequence.
- The hidden state lets later outputs depend on earlier inputs.
- A bidirectional variant can process a finite sequence in both directions.
doesn't solve
- RNN training can encounter vanishing and exploding gradients.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsRNN, PyTorch · read 27 Sept 2026
- docsBase RNN layer, Keras · read 27 Sept 2026
- paperLearning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation, Cho et al. · read 27 Sept 2026
- paperOn the difficulty of training Recurrent Neural Networks, Pascanu, Mikolov and Bengio · read 27 Sept 2026
- paperDeep learning, LeCun, Bengio and Hinton · read 27 Sept 2026