Matrices
A matrix is a rectangular grid of numbers with a fixed number of rows and columns. Machine learning keeps data and layer weights in matrices and multiplies them.
A matrix is a rectangular grid of numbers. Its shape is written as rows by columns, so a 2 × 3 matrix has 2 rows and 3 columns. A matrix with one row is called a row vector, and one with one column is a column vector. In machine learning code, a matrix is a tensor of rank 2: a single number has rank 0 and a vector has rank 1.
In a dataset, each row is usually one example and each column one feature. A greyscale image is a matrix too, indexed by row and column, with one number per pixel. The sample camera photo that comes with scikit-image has the shape 512 × 512. A colour image adds a third dimension for its colour channels. A batch is the set of examples used in one training step. Stacked into a matrix, a batch makes batch size one of the dimensions that the layer’s multiplication works over.
When two matrices are multiplied, each number in the product comes from one row of the first matrix and one column of the second, multiplied pair by pair and added up. That sum is the dot product of the row and the column. So two matrices can be multiplied only when the first one’s column count equals the second one’s row count. A 2 × 3 matrix times a 3 × 4 matrix gives a 2 × 4 result, and NumPy raises an error when the sizes do not match. Order matters, because AB and BA are generally not the same. It is also not the same as multiplying matching entries one by one, which is a separate operation called the Hadamard product.
In PyTorch, the Linear layer applies an affine transformation, y = xAᵀ + b, a matrix multiplication followed by adding a bias. It stores its weight matrix with the shape (out_features, in_features). NVIDIA’s documentation calls general matrix multiplication a fundamental building block for fully connected, recurrent and convolutional layers. NVIDIA’s Tensor Cores, introduced with the Volta GPU architecture, speed up matrix multiply-and-accumulate operations. Tensor Cores work on small blocks of a matrix at a time. Google’s TPUs contain matrix multiplication units, each a grid of 128 × 128 or 256 × 256 multiply-accumulators called a systolic array. Each multiply-accumulator passes its result straight to the next. As a result, no memory access is needed during the multiplication itself.
Machine learning needs one tidy way to hold many numbers and transform them all at once.
Follow one input through a layer's weight matrix.
- 1 · shapeA 2 × 3 weight matrix can multiply a column of 3 numbers, because the matrix's column count matches the input's length.
- 2 · pairTake one row of the matrix and the input column, and multiply them entry by entry.
- 3 · sumAdd those products to get one output number, the dot product of that row and the column.
- 4 · repeatDo the same for every row, so a matrix with 2 rows produces 2 outputs.
- 5 · biasA neural-network layer then adds a learned bias to each output.
Every output number is one row times one column. A layer repeats that for every row and every input.
| Who | What they ask | What it works with |
|---|---|---|
| Data scientist | “How do I turn this customer spreadsheet into something a model can read?” | A table with one row per customer and one column per feature |
| Computer vision engineer | “What does the model actually receive when I give it a photo?” | A grid of pixel values, one number per pixel for greyscale |
| Machine learning engineer | “Why will these two layers not connect?” | Weight matrix shapes whose inner sizes must match |
| Performance engineer | “Why is this layer slow on the GPU?” | The sizes of the matrix multiplications the layer runs |
- Holds a whole dataset or batch in one object with a known shape.
- Describes a whole layer as one operation, weights times inputs plus a bias.
- Lets hardware built for matrix operations run a model quickly.
- Catches wiring mistakes early, because mismatched shapes cannot be multiplied.
- A matrix holds numbers, so text, categories and images must first be turned into numbers.
- It does not choose the features; picking good columns is separate, hard work.
- Matrix multiplication is linear, so networks add activation functions to learn curved relationships.
- Speed is not automatic, because a multiplication that is too small can leave a GPU short of its peak.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMathematics for Machine Learning (book PDF), Deisenroth, Faisal and Ong, Cambridge University Press · read 27 Sept 2026
- docsnumpy.matmul, NumPy · read 27 Sept 2026
- docsLinear (PyTorch 2.14 documentation), PyTorch · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsA crash course on NumPy for images, scikit-image · read 27 Sept 2026
- docsLinear/Fully-Connected Layers User's Guide, NVIDIA · read 27 Sept 2026
- docsMatrix Multiplication Background User's Guide, NVIDIA · read 27 Sept 2026
- docsGPU Performance Background User's Guide, NVIDIA · read 27 Sept 2026
- docsTPU architecture, Google Cloud · read 27 Sept 2026