Concepts

Outliers

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

An outlier is a data point that sits an unusually long way from the rest. It may be a mistake to fix or a rare real event worth keeping.

1 · What it is

An outlier is a data point that sits an unusually long way from the rest. How far is too far is left to the analyst to decide. Some outliers are mistakes, such as a value typed in wrongly or an experiment that was not run properly. Google’s course gives the example of a person accidentally typing an extra digit. Others are real, caused by ordinary random variation or by something genuinely interesting. A large dataset will almost certainly contain a few rare but real extremes.

Outliers can even make a model’s weights overflow during training. Spotting them starts with simple checks. A z-score says how many standard deviations a value sits from the mean. In small samples z-scores can mislead, because the largest possible one is (n − 1) ÷ √n. With nine values that ceiling is 8 ÷ 3, about 2.67, so a beyond-3 rule could never fire. Iglewicz and Hoaglin instead recommend a modified z-score built from the median and the median absolute deviation. Values scoring above 3.5 are labelled potential outliers. For data with many features, model-based detectors such as Isolation Forest help. Isolation Forest splits the data at random, and unusual points need noticeably fewer splits to be cut off on their own. It works because anomalies tend to be few and different.

What to do next depends on the cause. A value shown to be wrong should be corrected, or deleted if it cannot be fixed. In Google’s data-centre example, readings of 1,000 degrees are mistakes to delete. Rare hot days between 31 and 45 degrees are real, and clipping them is reasonable. Deleting outliers blindly is risky, because they often carry useful information about the process or about how the data was recorded. Sometimes the outlier is the point, such as an unusual card payment that signals fraud. When the cause is unclear, the NIST handbook advises against simply deleting the value. Resistant tools help instead, such as scaling with the median and IQR. Huber loss gives far-off points a smaller pull without ignoring them. In scikit-learn’s comparison, a ridge regression line was strongly pulled by outliers while the Huber line was not.

2 · Why it exists

One far-off value can quietly change what the data seems to say.

Averages get draggedExtreme values pull the mean towards them, so it stops describing a typical value. The median barely moves.
Fitted lines tiltLeast squares, the usual way to fit a straight line, is very sensitive to unusual points. One or two outliers can seriously skew it.
Big misses dominateSquared error turns a miss of 3 into 9. In a five-example table, that one miss makes up about 56% of the mean squared error.
3 · How it works

Follow nine delivery times through the 1.5 × IQR rule.

Illustrative numbers. Only 95 lands beyond a fence, and taking it out moves the mean far more than the median.
  1. 1 · sortSort the values and find the quartiles, Q1 and Q3, which mark the 25th and 75th percentiles.
  2. 2 · spreadSubtract Q1 from Q3 to get the interquartile range (IQR), the width of the middle half of the data.
  3. 3 · fenceSet a lower fence at Q1 minus 1.5 IQRs and an upper fence at Q3 plus 1.5 IQRs.
  4. 4 · flagMark any value beyond a fence as a mild outlier, and any value beyond 3 IQRs as an extreme one.
  5. 5 · investigateBefore removing anything, find out why the value appeared.

A flag is a question, not a verdict. It marks a value for a closer look.

4 · Where it's used
WhoWhat they askWhat it works with
Delivery operations“Why does one route average 40 minutes when most trips take about 30?”Daily delivery times for each route
Data centre team“Should the readings of 1,000 degrees stay in the temperature feature?”Hourly sensor temperatures
Card payments team“Which transactions look nothing like this customer's usual spending?”Transaction amounts and merchants
Lab analyst“Is this one reading far above the others a recording slip?”Repeated measurements from one instrument
5 · What it solves, and what it doesn't
solves
  • Box plot fences give a simple, repeatable rule for spotting values worth a second look.
  • The median and IQR, which RobustScaler uses, are far less swayed by a few extreme values than the mean and standard deviation.
  • Clipping caps extreme values at a chosen limit, so a model stops overreacting to them.
  • Huber loss switches to absolute error for large misses, so outliers pull less on a fitted line.
doesn't solve
  • No rule says for certain what an outlier is. Where to draw the line is always a choice.
  • A flag does not say why a value is extreme, and sometimes nobody can tell whether it is bad data.
  • RobustScaler does not remove outliers. They are still in the scaled data.
  • Clipping makes a model treat every value above the cap alike, so a 45-degree day looks like a 35-degree one.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsWhat are outliers in the data?, NIST/SEMATECH e-Handbook of Statistical Methods · read 27 Sept 2026
  2. docsDetection of Outliers, NIST/SEMATECH e-Handbook of Statistical Methods · read 27 Sept 2026
  3. docsMeasures of Location, NIST/SEMATECH e-Handbook of Statistical Methods · read 27 Sept 2026
  4. docsLinear Least Squares Regression, NIST/SEMATECH e-Handbook of Statistical Methods · read 27 Sept 2026
  5. docsBox Plot, NIST/SEMATECH e-Handbook of Statistical Methods · read 27 Sept 2026
  6. docsNumerical data: Normalization (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  7. docsNumerical data: Scrubbing (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  8. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  9. docsNovelty and Outlier Detection, scikit-learn · read 27 Sept 2026
  10. docsRobustScaler, scikit-learn · read 27 Sept 2026
  11. docsCompare the effect of different scalers on data with outliers, scikit-learn · read 27 Sept 2026
  12. docsHuberRegressor, scikit-learn · read 27 Sept 2026
  13. docsHuberRegressor vs Ridge on dataset with strong outliers, scikit-learn · read 27 Sept 2026
  14. paperIsolation Forest, Liu, Ting and Zhou, IEEE ICDM 2008 · read 27 Sept 2026