Confusion matrix
A confusion matrix is a table that counts a classifier's predictions by what it predicted and what was really true, so you see which mistakes it makes.
A confusion matrix is a table of a classification model’s right and wrong predictions, with one cell for every pairing of predicted class and real class. A binary classifier chooses between two classes, a positive class it is testing for and a negative class. The positive class is the one the model is looking for, even when it is bad news, such as spam or a tumour. For a spam filter, a true positive is spam correctly sent to the spam folder. A false positive is a real email wrongly filed as spam. A false negative is spam that gets past the filter into the inbox. A true negative is a real email correctly left in the inbox. Every prediction lands in exactly one of those four cells.
Many models output a score, not a label. A classification threshold is the cut-off a person picks; scores above it become the positive class and scores below it the negative class. A score that lands exactly on the threshold is handled differently by different software; Keras predicts the negative class. In the made-up example in the figure, a threshold of 0.5 gives 2 true positives, 1 false positive, 1 false negative and 2 true negatives. Raise the threshold to 0.8 and only e1 is called spam, so the counts become 1 true positive, 0 false positives, 2 false negatives and 3 true negatives. A higher threshold means fewer false alarms but more missed spam.
With more than two classes, the table has one row and one column per class. Cells on the diagonal, where the real class and the predicted class match, are correct answers. Every other cell names one specific confusion. Google’s glossary shows a three-class model that sorts iris flowers. When the flower was really Virginica, the model got 109 right, called it Versicolor 27 times and called it Setosa only twice. A digit reader’s table might likewise show that it often writes 9 when the digit is 4.
Before reading any confusion matrix, check which way round it is drawn. scikit-learn puts the real class on the rows and the prediction on the columns. TensorFlow uses the same layout. Google’s crash course draws its spam table with the prediction on the rows instead. scikit-learn itself warns that other references may swap the axes.
The four counts are the raw material for most classification scores, such as accuracy, precision, recall and F1. Accuracy is the two correct cells divided by all four. A ROC curve repeats the count at many thresholds and plots two rates taken from the table. The true positive rate is the ROC curve’s y-axis. The false positive rate is its x-axis.
A single right-or-wrong total hides the mistakes that matter.
Follow six emails from score to table.
- 1 · scoreThe model gives each example a score between 0 and 1, such as how likely an email is to be spam.
- 2 · thresholdA threshold chosen by a person turns each score into a prediction, yes above it and no below it.
- 3 · pairEach prediction is put next to the example's real answer, its ground truth.
- 4 · sortThe pair picks one cell: predicted yes and really yes is a true positive, yes but really no is a false positive, no but really yes is a false negative, and no and really no is a true negative.
- 5 · countThe cells are added up, and those four counts are the inputs for the usual classification scores.
The table does not grade the model with one number. It keeps false alarms and misses apart.
| Who | What they ask | What it works with |
|---|---|---|
| Email spam team | “How many real emails end up in the spam folder?” | The false-positive cell at the current threshold |
| Security team | “How many harmless websites did the malware filter block?” | Legitimate sites the model predicted as malware |
| Handwriting recognition team | “Which digits does our reader mix up most often?” | The off-diagonal cells of a ten-by-ten digit table |
| Medical screening team | “How many sick patients did the test miss?” | The false-negative cell against confirmed diagnoses |
- It shows what kind of mistake a model makes, not only how many.
- With many classes, it shows which pairs of classes the model confuses.
- Its row totals count the real examples in each class, which shows whether the classes are balanced.
- Its counts are enough to calculate accuracy, precision, recall and related scores.
- It describes one threshold only. Move the threshold and the four counts change.
- It does not say which mistake costs more. People have to decide that.
- It cannot be more correct than the ground truth it is counted against.
- Four counts are hard to compare at a glance, and squeezing them into one score can mislead.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsClassification: Thresholds and the confusion matrix (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docssklearn.metrics.confusion_matrix, scikit-learn · read 27 Sept 2026
- docs3.4. Metrics and scoring: quantifying the quality of predictions, scikit-learn · read 27 Sept 2026
- docstf.math.confusion_matrix, TensorFlow · read 27 Sept 2026
- paperGlossary of Terms (Machine Learning 30, 271–274, 1998), Kohavi and Provost, Machine Learning journal · read 27 Sept 2026
- paperThe advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation, Chicco and Jurman, BMC Genomics 2020 · read 27 Sept 2026