Precision and recall
Precision asks how many of the things a model flagged were right. Recall asks how many of the real cases it managed to find.
Precision and recall are two scores for a model that sorts things into yes and no. Precision asks: when the model said yes, how often was it right? It is the true positives divided by the true positives plus the false positives. Recall asks: of all the real yes cases, how many did the model find? It is the true positives divided by the true positives plus the false negatives. A false positive is a false alarm, a harmless case wrongly called positive. A false negative is a real case the model missed. Recall is also called sensitivity or the true positive rate. Precision is also called positive predictive value.
Take an illustrative spam filter checking 100 emails, 25 of them spam. It flags 20, and 15 of those really are spam, so precision is 15 ÷ 20 = 0.75. Recall is 15 ÷ 25 = 0.60, because 10 spam emails slipped into the inbox. The 70 true negatives appear in neither formula. In information retrieval, the field that studies search engines, precision and recall are the two most common and basic ways to judge results. There, precision asks how much of what came back was useful, and recall asks how much of the useful material came back. A typical web searcher wants every hit on page one to be relevant, which means high precision. Paralegals and intelligence analysts push for the highest recall they can get and put up with fairly weak precision to reach it. A search that hands back every document it has scores perfect recall, but its precision collapses.
Move the threshold up and the filter gets pickier: fewer innocent emails get caught, but more spam slips through. That usually pushes precision up and recall down, so the two tend to pull against each other. At a threshold of 0.9, the illustrative filter might flag 10 emails, 9 of them spam: precision 0.90, recall 9 ÷ 25 = 0.36. The bottom of the recall fraction stays at 25 whatever the threshold; only the count of spam caught above the line changes. This is a tendency, not a law. Lowering the threshold can leave recall where it was while precision bounces up or down.
No score is always the one to chase. It comes down to what each kind of mistake would cost in your situation. Google’s crash course says to lean on recall when a missed case hurts more than a false alarm. In disease prediction, missing someone who is ill typically does more damage than a false alarm. For a spam filter, Google says any of three choices can be sensible: favour recall to catch every spam email, favour precision so the spam folder holds only spam, or strike a balance. Plotting precision against recall at every threshold gives a precision-recall curve. For a ranked list of search results, the curve has a saw-tooth shape. At its low-threshold end, a classifier that flags everything has recall 1 and precision equal to the share of positives. A 2015 study argued that on very imbalanced data this curve is more informative than the better-known ROC plot, because precision looks only at how the model’s positive calls turned out.
One score cannot describe two different kinds of mistake.
Follow one batch of emails through a spam filter.
- 1 · scoreThe model gives each email a probability, between 0 and 1, that it is spam.
- 2 · cutEmails scoring above a chosen threshold are flagged as spam, and the rest are not.
- 3 · countEach decision is checked against the truth and counted as a true positive, false positive, false negative or true negative.
- 4 · dividePrecision divides the true positives by everything flagged, while recall divides them by every email that really was spam.
Same top number, different bottoms: precision divides by what the model flagged, recall by what was really there.
| Who | What they ask | What it works with |
|---|---|---|
| Email provider | “How much of the spam folder is really spam?” | Emails the filter flagged, scored for precision |
| Disease screening team | “How many sick patients did the test miss?” | Every confirmed case, scored for recall |
| Legal research team | “Did our search turn up every relevant document?” | All documents relevant to the case, scored for recall |
| Web search team | “Is every result on the first page relevant?” | The top results for a query, scored for precision |
- Keeps false alarms and misses apart, while accuracy folds every right answer, positive or negative, into one number.
- Stays meaningful when the thing you are looking for is rare, where accuracy can mislead.
- Gives a way to choose a threshold by weighing the costs of each kind of mistake.
- Focuses on the true positives, so a huge pile of correctly ignored cases cannot inflate the score.
- Neither number is enough alone. Flagging everything gives perfect recall and very low precision.
- Precision stops working when a model flags nothing, because zero divided by zero is not a number.
- Neither score counts true negatives, so neither shows how many harmless cases were correctly left alone.
- The baseline of a precision-recall curve moves with how rare the positives are, so curves from differently balanced data are hard to compare.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsClassification: Accuracy, recall, precision, and related metrics (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsThresholds and the confusion matrix (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsprecision_recall_curve, scikit-learn · read 27 Sept 2026
- docsPrecision-Recall, scikit-learn · read 27 Sept 2026
- paperEvaluation of unranked retrieval sets (Introduction to Information Retrieval, chapter 8), Manning, Raghavan and Schütze, Cambridge University Press · read 27 Sept 2026
- paperEvaluation of ranked retrieval results (Introduction to Information Retrieval, chapter 8), Manning, Raghavan and Schütze, Cambridge University Press · read 27 Sept 2026
- paperThe Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets, Saito and Rehmsmeier, PLOS ONE 2015 · read 27 Sept 2026