Concepts

Anomaly detection

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Anomaly detection finds the few items or events that do not fit the usual pattern in data, such as a fraudulent card payment or an odd burst of network traffic.

1 · What it is

An odd card payment can be a sign of fraud. An unusual pattern of network traffic can point to someone getting in without permission. A sudden rise in a product’s return rate can point to a defect or to fraud. There are three settings. In outlier detection, some anomalies are already mixed into the training data. The detector learns where points pile up and treats the scattered few as suspects. scikit-learn calls this unsupervised anomaly detection. Novelty detection trains on clean data, then checks each new arrival against that picture of normal. scikit-learn calls that semi-supervised anomaly detection. When labelled anomalies already exist, an ordinary supervised model such as boosted trees or a random forest can learn to predict them.

The methods fall into a few families. Statistical methods assume a known shape for normal data, such as a Gaussian, and flag points that sit far from it. Local Outlier Factor, or LOF, compares how crowded a point’s neighbourhood is with how crowded its neighbours’ neighbourhoods are. A One-Class SVM learns a boundary around clean training data and treats points outside it as abnormal. Isolation methods cut the data at random and watch which points get separated quickly. Amazon’s Random Cut Forest uses a related idea. Each tree’s score roughly tracks how shallow the point lands, and the forest averages those scores. Autoencoder and PCA models flag rows they rebuild badly, measured as reconstruction error. For time series, a forecasting model predicts a range for each moment and flags values that fall outside it.

Most detectors produce a score first and a yes or no second. For Random Cut Forest, AWS notes a common rule of flagging scores more than three standard deviations above the mean score. In scikit-learn, a contamination setting, the expected share of outliers, places the threshold instead. Because anomalies are few, judge the result with precision and recall, not accuracy. Precision is the share of flagged cases that were real anomalies. Recall is the share of real anomalies that were flagged. Raising the threshold usually cuts false alarms but misses more real cases, so the two pull against each other. Checking a live detector is hard, because real-world data is often unlabelled.

2 · Why it exists

Anomalies are few, hard to define in advance and may come without labels.

No clear definitionIt can be hard to say ahead of time what counts as anomalous. Often there are no labelled examples to train a model on either.
Accuracy misleadsIf only 1% of cases are positive, a detector that never flags anything still reaches 99% accuracy, yet it catches nothing.
Too many false alarmsA detector tuned only to describe normal data can flag too many normal cases or miss real anomalies.
3 · How it works

Follow one odd point through an isolation forest.

Illustrative points. The odd point is cut off after 2 random cuts, the normal one after 6. Averaged over many trees, a short path means a high anomaly score.
  1. 1 · sampleEach tree is grown on a small random sample of the data, 256 points by default in the original paper.
  2. 2 · cutThe tree picks a random feature and a random split value between that feature's minimum and maximum, over and over.
  3. 3 · countThe number of cuts needed to leave a point on its own is its path length, and anomalies get noticeably shorter paths.
  4. 4 · scorePath lengths are averaged over many trees and turned into a score, where values very close to 1 mark clear anomalies.
  5. 5 · alertA threshold on that score, for example set from the expected share of outliers, decides which points become alerts.

Anomalies are few and different, so random cuts reach them early.

4 · Where it's used
WhoWhat they askWhat it works with
Card payments team“Which of today's transactions look nothing like normal spending?”Transaction amounts, merchants and times
Network security team“Is this burst of traffic a sign of someone breaking in?”Connection logs for each server
Online shop“Why did returns for this kettle jump this week?”Daily return rates for each product
Operations team“Does this jump in a metric point to a technical problem?”The metric's values over time
5 · What it solves, and what it doesn't
solves
  • It can find unusual points without labelled examples of anomalies.
  • It gives every row a score, so a team can rank the data and review the oddest cases first.
  • Isolation Forest runs in linear time with low memory, so it scales to large datasets.
  • For time series, a forecast gives lower and upper bounds, and values outside them are flagged.
doesn't solve
  • Unsupervised outlier detection assumes anomalies sit in sparse regions, so a dense cluster of anomalies can slip through.
  • Many anomalies can hide each other, which is called masking. Normal points near anomalies can be wrongly flagged, which is called swamping.
  • The threshold is a choice. What counts as a low or high score depends on the application.
  • A detector left alone goes stale. If live data changes and the model is not retrained, its quality declines.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsNovelty and Outlier Detection (scikit-learn user guide), scikit-learn · read 27 Sept 2026
  2. paperIsolation Forest, Liu, Ting and Zhou, IEEE ICDM 2008 · read 27 Sept 2026
  3. docsIsolationForest (scikit-learn API reference), scikit-learn · read 27 Sept 2026
  4. docsRandom Cut Forest (RCF) Algorithm, Amazon Web Services · read 27 Sept 2026
  5. docsHow RCF Works, Amazon Web Services · read 27 Sept 2026
  6. docsAnomaly detection overview, Google Cloud · read 27 Sept 2026
  7. docsThe ML.DETECT_ANOMALIES function, Google Cloud · read 27 Sept 2026
  8. docsClassification: Accuracy, recall, precision, and related metrics (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  9. docsProduction ML systems: Monitoring pipelines (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026