Concepts

Overfitting

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Overfitting happens when a model learns its training examples so specifically that it performs poorly on new examples.

1 · What it is

A flexible model can discover real signal, but it can also fit random quirks in a finite training sample. Those quirks exist in the training set but not in fresh data, so fitting them buys nothing later. The result is overfitting: excellent memory of the sample and weak generalization.

Checking on unseen validation data tests whether the learned rule transfers beyond that sample. When training and validation loss drift apart, treat that gap as a warning sign. Leakage between the splits can make that comparison look better than it really is.

Overfitting can come from unrepresentative data, an overly complex model, or both. Common fixes include ending training once validation results begin to slip and adding a penalty for large weights. Keep the final test set out of every one of those choices.

2 · Why it exists

A low training error can hide a model that learned noise or accidental details instead of a reusable pattern.

False confidenceAn overfit model can score extremely well on examples it already saw.
Poor transferThe same model can make poor predictions on new data.
Leaky checksIf test rows sneak into a preprocessing step such as feature selection, even the reported holdout score can look too good.
3 · How it works

Follow one shortcut from training success to a new-data failure.

A returns model memorizes an accidental tracking-code shortcut Training orders contain an accidental tracking-code pattern. A flexible model memorizes it, scores perfectly on training rows, then fails on a new order where that coincidence does not hold. TRAINING ORDERS tracking · item · label R7… · shoes · return K2… · lamp · keep R7… · coat · return R7 is only a coincidence KEY STEP Flexible fit memorizes shortcut R7 → return training: 3 / 3 correct NEW ORDER R7… · lamp · keep prediction: return wrong on unseen data shortcut does not survive the split Training success: excellent Generalization: poor
The model fits an accidental tracking-code shortcut, so its training success does not transfer.
  1. 1 · sampleThe training set contains both a useful pattern and coincidences that will not persist.
  2. 2 · fitA more flexible model has room to fit those coincidences as well as the real pattern.
  3. 3 · compareTraining loss keeps falling or levels off, while validation loss turns upward.
  4. 4 · correctCharging a cost for big weights nudges the model toward a simpler rule, leaving less room to overfit.

Overfitting is a gap between remembering the sample and generalizing beyond it.

4 · Where it's used
WhoWhat they askWhat it works with
Fraud team“Did the model learn fraud signals or yesterday's merchant IDs?”Training and validation scores by time period
Vision engineer“Is the classifier using the object or a camera watermark?”Errors on images from unseen cameras
Model trainer“Did another epoch help outside the training set?”Training and validation loss curves
5 · What it solves, and what it doesn't
solves
  • Naming overfitting separates training-set fit from performance on unseen examples.
  • Validation curves can expose it when training performance stays strong but validation performance weakens.
  • Once named, it points to fixes such as early stopping, which halts training before the model has finished converging.
doesn't solve
  • A good validation score only carries over when real-world data looks statistically like the training and validation splits.
  • Blaming model size alone misses the other common cause, which is training data that fails to reflect real use.
  • Cross-validation scores can still come out too rosy when held-out data leaks into the model-building steps.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsOverfitting, Google for Developers · read 27 Sept 2026
  2. docsUnderfitting vs. Overfitting, scikit-learn · read 27 Sept 2026
  3. docs3.5. Validation curves: plotting scores to evaluate models, scikit-learn · read 27 Sept 2026
  4. paperDropout: A Simple Way to Prevent Neural Networks from Overfitting, Journal of Machine Learning Research · read 27 Sept 2026
  5. docs12. Common pitfalls and recommended practices, scikit-learn · read 27 Sept 2026
  6. docsOverfitting: L2 regularization, Google for Developers · read 27 Sept 2026