Cosine similarity
Cosine similarity scores how closely two vectors point the same way, ignoring how long they are. The score runs from −1 to 1.
Cosine similarity is a score for how closely two vectors point the same way. To get it, take the dot product of the two vectors and divide it by their two lengths multiplied together. The result is the cosine of the angle between the two arrows. Vectors pointing the same way score 1, vectors at right angles score 0, and vectors pointing in opposite directions score −1. Length drops out of the answer. Make one arrow four times longer and it still points the same way, so its angle to any other arrow, and its score, stay exactly as they were.
It is a standard tool in information retrieval, the study of how search engines find documents. The classic vector space model turns each document into a vector with one position per word. Plain distance between two such vectors misleads you: a long article and a short summary of the same story end up far apart, just because the long one piles up bigger word counts. Dividing by the lengths cancels that out. That is why cosine became the usual score for comparing documents and for “more like this” links. Embeddings inherited the habit. In OpenAI’s retrieval example, each document is scored against the question this way and the best scores win. Google’s Gemini docs pick cosine for the same reason: two texts about one idea should count as close whatever their vector lengths.
Many embedding models already return vectors scaled to length 1. OpenAI’s embeddings are, and so are Gemini’s default embeddings of 3,072 numbers. With both lengths equal to 1, the division changes nothing, so cosine similarity equals the plain dot product and ranks results the same way as straight-line distance. Vector tools use that shortcut: pgvector has a cosine distance operator but advises inner product for normalised vectors, and Faiss runs cosine search by normalising vectors first. Cosine distance, which some tools report instead, is simply 1 minus the similarity.
Comparing vectors by plain distance or a raw dot product lets length get in the way.
Follow two pairs of small vectors through the calculation.
- 1 · multiplyMultiply the two vectors position by position and add the results, which gives their dot product.
- 2 · measureFind each vector's length, its straight-line distance from the origin.
- 3 · divideDivide the dot product by the two lengths multiplied together, which removes the effect of how long each vector is.
- 4 · readRead the result as a score where 1 means the same direction, 0 means at right angles and −1 means opposite.
Stretching a vector changes its length but not its direction, so its cosine similarity with anything stays the same.
| Who | What they ask | What it works with |
|---|---|---|
| Search team | “Which help articles best match this question?” | Scores between the question's embedding and each article's embedding |
| Recommendation team | “Which stories are most like the one this reader just finished?” | Embedding vectors of every story in the archive |
| Database engineer | “How do I sort rows by closeness to a query vector in Postgres?” | The pgvector cosine distance operator on an embedding column |
| Developer tools team | “Which function in this repository does this plain-English request describe?” | Embeddings of every function in the codebase |
- Compares direction only, so a long document and a short one on the same topic can still score as close.
- Gives every pair a score on the same fixed scale, from −1 to 1.
- Ranks stored embeddings against a query for semantic search, retrieval and recommendation.
- On vectors already scaled to length 1, it becomes a plain dot product, which is quicker to compute.
- It ignores length, so anything stored in a vector's size, such as how popular an item is, is lost.
- Its scores depend on how the embedding model was trained, and for some models they can be arbitrary.
- An all-zero vector has no direction, so code has to guard against dividing by zero.
- It only sees what the vectors hold, and word-count vectors cannot tell "Mary is quicker than John" from the reverse.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docs8.8. Pairwise metrics, Affinities and Kernels (User Guide), scikit-learn · read 27 Sept 2026
- docscosine_distances, scikit-learn · read 27 Sept 2026
- docsCosineSimilarity, PyTorch · read 27 Sept 2026
- paperMathematics for Machine Learning (book PDF), chapter 3, Cambridge University Press (Deisenroth, Faisal and Ong) · read 27 Sept 2026
- paperDot products (Introduction to Information Retrieval, chapter 6), Cambridge University Press (Manning, Raghavan and Schütze) · read 27 Sept 2026
- docsMeasuring similarity from embeddings, Google for Developers · read 27 Sept 2026
- docsVector embeddings, OpenAI · read 27 Sept 2026
- docsEmbeddings (Gemini API), Google AI for Developers · read 27 Sept 2026
- repopgvector: Open-source vector similarity search for Postgres, pgvector · read 27 Sept 2026
- repoMetricType and distances (Faiss wiki), Meta (Faiss) · read 27 Sept 2026
- paperIs Cosine-Similarity of Embeddings Really About Similarity?, Steck, Ekanadham and Kallus, arXiv 2024 · read 27 Sept 2026