Concepts
Convolutional neural networks
1 · In one line
A convolutional neural network uses sparse convolutions that reuse the same weights at multiple locations in an ordered grid.
1 · What it is
A convolutional layer applies a 2D convolution over an input signal. Each output value sums products between a kernel and the input elements it overlaps. The same kernel weights are used at every position reached by the sliding operation.
Stride is the distance between consecutive kernel positions. Zero padding concatenates zeros at the beginning and end of an axis.
Multiple processing layers can learn representations with multiple levels of abstraction.
Images have width and height axes for which ordering matters.
Grid structureFlattening an image into one vector removes its explicit width-and-height layout.
Position changesA detector should apply the same rule when a pattern appears at a different position.
Local evidenceOnly a few input units contribute to one output unit in a sparse convolution.
Follow one learned filter across a small image.
- 1 · windowSelect a small local window from the input grid.
- 2 · matchMultiply each kernel element by the input element it overlaps.
- 3 · writeSum those products to obtain the output at the current location.
- 4 · slideMove the kernel by its stride distance to the next position.
- 5 · stackRepeat the procedure with different kernels to form multiple output feature maps.
The defining move is weight sharing: one filter is reused across positions.
| Who | What they ask | What it works with |
|---|---|---|
| Medical-imaging team | “Which regions contain a suspicious texture?” | Local patterns in scan pixels |
| Factory inspector | “Does this component contain a visible defect?” | Camera frames from the production line |
| Audio engineer | “Where does this sound pattern occur?” | Local regions in a time-frequency grid |
solves
- Convolution is sparse and reuses the same parameters at multiple input locations.
- A convolution produces output feature maps.
- Kernel size, stride, padding and dilation provide explicit control over receptive geometry and output shape.
doesn't solve
- A 3 by 3 filter has a 3 by 3 receptive field at that layer.
- Deeper neural networks are more difficult to train.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsConv2d, PyTorch · read 27 Sept 2026
- paperA guide to convolution arithmetic for deep learning, Dumoulin and Visin · read 27 Sept 2026
- paperDeep learning, LeCun, Bengio and Hinton · read 27 Sept 2026
- paperVery Deep Convolutional Networks for Large-Scale Image Recognition, Simonyan and Zisserman · read 27 Sept 2026
- paperDeep Residual Learning for Image Recognition, He, Zhang, Ren and Sun · read 27 Sept 2026