Concepts

Accuracy

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Accuracy is the share of a model's predictions that were right, the correct answers divided by all answers.

1 · What it is

In machine learning, accuracy is the share of predictions a model got right: the number correct divided by the total. Get 40 of 50 test cases right and your accuracy is 80%. For a yes-or-no model, the right answers are true positives (a correct yes) and true negatives (a correct no). The wrong ones are false positives, where the model says yes to a real no, and false negatives, where it misses a real yes. So accuracy is (TP + TN) ÷ (TP + TN + FP + FN). The same four counts, laid out as a confusion matrix, also give related scores such as precision, recall and the F1 score, which combines those two.

The everyday word is looser. Machine learning ties accuracy, precision and recall to precise formulas, which are narrower than how people use those words day to day. In medicine, a new test’s diagnostic accuracy is about how closely its results line up with an accepted reference standard, and the US Food and Drug Administration (FDA) describes it with pairs of numbers such as sensitivity and specificity. Language model benchmarks report plain accuracy too. MMLU is a test of multiple-choice questions covering 57 tasks, and it measures the share a model answers correctly. On MMLU, guessing at random scores 25%.

Accuracy misleads most on imbalanced data, where one class is far more common than the other. Such data is the norm: among card payments, fraud can be fewer than 1 in 1,000. Google’s glossary imagines a city where it snows on 25 days per century. A model that predicts no snow every day scores 99.93%, yet it tells you nothing about when snow will come. In general, a model that always picks the larger class scores that class’s share of the test set. Balanced accuracy, built to avoid inflated scores on imbalanced data, averages the recall of each class, meaning the share of that class’s real cases the model got right. If a high score came only from the imbalance, balanced accuracy falls to chance. On a balanced test set it equals ordinary accuracy.

2 · Why it exists

One percentage can hide what a model actually gets wrong.

Rare cases vanishIf just 1 case in 100 is a yes, a lazy model that says no to everything is right 99 times out of 100 and never catches a single yes.
Every mistake weighs the sameFlagging a healthy patient and missing a sick one both add one to the error count, though in real life the two mistakes rarely cost the same.
It moves with the thresholdMost models output a score, and someone picks the line that turns it into yes or no. Slide that line and the accuracy figure slides with it.
3 · How it works

Follow one test set through the calculation.

Illustrative numbers. A model that answers no every time wins on accuracy while catching none of the cases that matter.
  1. 1 · predictThe model scores each test case, and a threshold chosen by a person turns each score into yes or no.
  2. 2 · compareEach answer is checked against the true label and named a true or false positive or negative.
  3. 3 · divideAccuracy is the correct answers, true positives plus true negatives, divided by all answers.
  4. 4 · checkThe result is compared with a model that always predicts the most common class, which scores that class's share for free.

Accuracy counts how many answers were right, not which ones.

4 · Where it's used
WhoWhat they askWhat it works with
Student building a digit reader“Did the new version beat the old one on a balanced test set?”Share of test images given the right digit
Card fraud team“Our model is 99.9% accurate, so why does fraud still get through?”Accuracy set beside the count of missed fraud cases
Hospital AI team“Is our high score just the healthy majority talking?”Accuracy compared with balanced accuracy on scan results
Language model evaluators“What share of MMLU questions does the new model get right?”Chosen answers compared with the answer key
5 · What it solves, and what it doesn't
solves
  • Gives one simple number for how often a model is right, using all four outcomes.
  • When yes and no cases are about equal in number, it gives a fair first read on how good a model is.
  • It is a common default when no task-specific metric has been chosen.
  • It is one of the most widely used scores for yes-or-no classifiers.
doesn't solve
  • It can look excellent while missing every rare case, because always choosing the larger class scores that class's share.
  • It counts a false alarm and a miss the same way, as one wrong answer.
  • On lopsided data it flatters a model, which is why some researchers argue for the Matthews correlation coefficient.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsClassification: Accuracy, recall, precision, and related metrics (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  2. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  3. docsThresholds and the confusion matrix (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  4. docsDatasets: Class-imbalanced datasets (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  5. docs3.4. Metrics and scoring: quantifying the quality of predictions, scikit-learn developers · read 27 Sept 2026
  6. paperThe advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation, Chicco and Jurman, BMC Genomics 2020 · read 27 Sept 2026
  7. officialStatistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests, U.S. Food and Drug Administration · read 27 Sept 2026
  8. paperMeasuring Massive Multitask Language Understanding, Hendrycks et al., ICLR 2021 · read 27 Sept 2026