ROC curveConcepts

Receiver operating characteristic curve

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

An ROC curve shows, for every threshold a classifier could use, how many real positives it catches against how many false alarms it raises.

1 · What it is

An ROC curve is a graph of how a binary classifier behaves across all of its possible thresholds. A binary classifier sorts things into two groups, such as spam and not spam. It gives each item a score, and a threshold decides which group the item lands in. The vertical axis is the true positive rate, another name for recall: the share of real positives the model catches. The horizontal axis is the false positive rate: the share of real negatives it wrongly flags. The name, receiver operating characteristic, is left over from Second World War radar, where it helped tell a real signal from noise.

Take the made-up example in the figure: five real positives and five real negatives, each with a score. A threshold above every score predicts nothing positive, so both rates are 0. Lower it to 0.9 and one positive clears it, giving a true positive rate of 1 in 5, or 0.2, with no false alarms. At 0.5, four positives and two negatives clear it, so the point is 0.4 across and 0.8 up. At 0.3 it is 0.6 across and 1.0 up. Lowering the threshold only lets more examples through, so the points move up and to the right, joined by straight lines.

The area under that curve, AUC, has a neat meaning. It is the chance that a randomly chosen positive gets a higher score than a randomly chosen negative. Here, 20 of the 25 positive and negative pairs are ranked the right way round, so the full curve, using every threshold, has an AUC of 20 ÷ 25 = 0.80. The three thresholds in the figure give 0.78, a little less, because a coarser curve cuts a corner. A model that guesses at random follows the diagonal and scores 0.5; a perfect model scores 1.0. In the Python library scikit-learn, roc_curve returns the points and roc_auc_score returns the area.

2 · Why it exists

A classifier's yes-or-no answers depend on a threshold, and one score at one threshold hides the rest.

One setting onlyScores such as accuracy and recall are worked out at a single threshold. They cannot show how the model does across the other settings.
A human picks itThe threshold is chosen by a person, not learned in training. That choice strongly changes how many false positives and false negatives you get.
No obvious answerIt is often unclear in advance which threshold is the right one for a job.
3 · How it works

Follow ten scored examples as the threshold slides from high to low.

Made-up scores. Each threshold gives one point; straight lines join the points, and the area under them is the AUC.
  1. 1 · scoreThe model gives each example a score, such as its estimated probability of being positive.
  2. 2 · cutA threshold is chosen, and every example scoring at or above it is predicted positive.
  3. 3 · countThe true positive rate is the share of real positives caught, and the false positive rate is the share of real negatives wrongly flagged.
  4. 4 · plotEach threshold becomes one point, false positive rate across and true positive rate up, and sliding the threshold down traces the curve.
  5. 5 · summarizeThe area under the curve, or AUC, turns the whole curve into one number between 0 and 1.

Up and to the left is better: the top-left corner means every positive caught and no false alarms.

4 · Where it's used
WhoWhat they askWhat it works with
Email team“Where should the spam cut-off sit so that business mail is not lost?”Spam scores for a set of emails already labelled spam or not
Hospital laboratory“Which blood test value best separates patients with the disease from those without?”Test results for patients whose diagnosis is known
Fraud team“Which of our two models ranks real fraud above normal payments more often?”Both models' scores on the same held-out transactions
5 · What it solves, and what it doesn't
solves
  • It shows the trade-off at every possible threshold in one picture, instead of one setting at a time.
  • AUC gives one number for comparing two models, and on roughly balanced data the larger area is generally the better model.
  • It helps you pick a threshold. If false alarms are costly, choose a point with a low false positive rate. If misses are costly, choose one with a higher true positive rate.
  • Each axis is a rate inside one class, so the curve is not affected by how common the positive class is.
doesn't solve
  • It does not pick the threshold for you, and the threshold values are not shown on the curve itself.
  • On heavily imbalanced data the picture can be deceptive. A precision-recall curve, which tracks how many positive predictions are right, may be more informative there.
  • One AUC number hides detail. It can mislead when two curves cross, and it can miss differences among the highest-scoring examples.
  • It needs scores from the model. A classifier that only outputs yes or no gives just a single point.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsClassification: ROC and AUC, Google for Developers, Machine Learning Crash Course · read 27 Sept 2026
  2. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  3. docsroc_curve, scikit-learn · read 27 Sept 2026
  4. docsMetrics and scoring: quantifying the quality of predictions, scikit-learn · read 27 Sept 2026
  5. docsMulticlass Receiver Operating Characteristic (ROC), scikit-learn · read 27 Sept 2026
  6. paperThe Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets, Saito and Rehmsmeier, PLOS ONE (2015) · read 27 Sept 2026
  7. paperReceiver operating characteristic curve: overview and practical use for clinicians, Nahm, Korean Journal of Anesthesiology (2022), via PubMed Central · read 27 Sept 2026
  8. paperReceiver Operating Characteristic (ROC) Curve Analysis for Medical Diagnostic Test Evaluation, Hajian-Tilaki, Caspian Journal of Internal Medicine (2013), via PubMed Central · read 27 Sept 2026