Dot product
The dot product multiplies two equal-length lists of numbers position by position and adds the results into one number set by the angle between them and their lengths.
The dot product takes two lists of numbers of the same length and returns one number. You multiply the first entries together, then the second entries, and so on, and add up all the products. NumPy’s dot function gives this inner product when both inputs are one-dimensional arrays. PyTorch’s torch.dot is stricter and refuses anything except two flat, one-dimensional tensors of equal length.
The same number has a picture. It equals the length of one vector times the length of the other times the cosine of the angle between them, written a · b = ‖a‖‖b‖ cos θ. The dot product is positive when the vectors are less than a right angle apart and point roughly the same way. It is negative when they are more than a right angle apart and point roughly opposite ways. It is zero exactly when they are perpendicular, which mathematicians call orthogonal.
Length sneaks in, so of two arrows aimed the same way, the longer one earns the bigger score. In recommendations, a video that turns up again and again in training often gets a long embedding, so the plain dot product favours what is already popular. Divide the result by both lengths and they cancel out, leaving just the cosine of the angle, which is cosine similarity. Stretch or shrink every vector to length 1 before you start, and the three usual yardsticks (dot product, cosine and straight-line distance) end up telling you the same story.
Inside a neural network, a single neuron scales each input by a learned weight and totals the lot. That is the multiply-then-add recipe again, so the total is the dot product of the input list and the weight list. Multiplying two matrices repeats the trick, since every cell of the answer comes from pairing one row of the first matrix with one column of the second. The Transformer uses it to decide where to look; each query is dotted with every key, each score is divided by the square root of the key size, and softmax turns the scores into weights. The authors called this scaled dot-product attention. They added the division because, they suspected, keys with many entries give huge scores that push softmax into a zone of near-zero gradients. Faiss, a vector search library, compares vectors in two main ways, by L2 distance or by inner product. Faiss says its inner-product mode is typically used in recommendation systems, to find the stored items that score highest against a query. If cosine similarity is what you really want, Faiss’s advice is to normalise the vectors first and then run the same inner-product search.
Machine learning keeps having to turn two lists of numbers into one.
Follow two small vectors to one number.
- 1 · lineLine up two vectors with the same number of entries, such as a = (3, −1, 2) and b = (1, 4, 2).
- 2 · multiplyMultiply the entries that sit in the same position, giving 3, −4 and 4.
- 3 · addAdd those products to get a single number, here 3 − 4 + 4 = 3.
- 4 · readRead the sign, where positive means less than a right angle apart, zero means perpendicular and negative means roughly opposite.
Multiply, then add. The sign follows the angle, and the size also grows with the lengths.
| Who | What they ask | What it works with |
|---|---|---|
| Neural network engineer | “What does one neuron actually compute from its inputs?” | Input values and their weights, multiplied and added |
| Language model researcher | “How strongly should this word attend to each earlier word?” | One query vector against every key vector |
| Search engineer | “Which stored items score highest for this query?” | Embedding vectors in an inner-product index |
| Recommendation team | “Should popular videos get a boost for similar users?” | Item embeddings compared by dot product or cosine |
- Turns two vectors into one number with a clear geometric meaning.
- Gives a simple test for perpendicular vectors, because their dot product is zero.
- Runs as fast matrix multiplication, which is why dot-product attention is quick in practice.
- Supports maximum inner product search, which vector libraries such as Faiss offer.
- It grows with vector length, so popular items with long embeddings can skew it.
- On its own it is a different score from cosine similarity; the two only agree once every vector has length 1.
- It needs two vectors with the same number of entries.
- Keys with many entries can give huge scores, so the original Transformer divides by the square root of the key size.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMathematics for Machine Learning (book PDF), Deisenroth, Faisal and Ong, Cambridge University Press · read 27 Sept 2026
- docs18.02SC Notes: Dot Product, MIT OpenCourseWare · read 27 Sept 2026
- docs18.02 Lecture Notes, Fall 2021, MIT Department of Mathematics · read 27 Sept 2026
- docsMeasuring similarity from embeddings, Google for Developers · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- paperAttention Is All You Need, Vaswani et al., NeurIPS 2017 · read 27 Sept 2026
- repoMetricType and distances, Faiss (Meta) · read 27 Sept 2026
- docsnumpy.dot, NumPy · read 27 Sept 2026
- docstorch.dot (PyTorch 2.14 documentation), PyTorch · read 27 Sept 2026