Concepts

Data augmentation

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Data augmentation makes extra training examples by changing existing ones in small ways that keep the answer the same, such as flipping or cropping a photo.

1 · What it is

Data augmentation creates extra training examples by transforming ones you already have. A photo can be rotated, stretched or mirrored to make many variants, and each variant keeps the original label. The goal is more variety as well as more examples. A model should learn, for instance, that an object’s identity does not change when the lighting does.

The 2012 AlexNet image classifier set the standard recipe, which was still in use on ImageNet years later with small changes. Its authors cut random 224 by 224 pixel patches, and their mirror images, out of 256 by 256 photos. That enlarged the training set by a factor of 2048, though the copies were highly interdependent. A second trick shifted the colour of each photo and cut the top-1 error rate by over 1%. The new images were made on the CPU while the GPU trained on the previous batch, so they cost almost nothing. At test time AlexNet did something different, averaging its predictions over ten crops and mirror images of each test photo.

The idea reaches beyond pictures. One text method translates a sentence into French and back into English. Another method, EDA, replaces words with random synonyms. It can also swap the positions of two words. For speech, SpecAugment edits the spectrogram, a picture of how the sound changes over time. It masks out blocks of frequency channels and blocks of time steps, and warps the features in time. Older audio methods sped recordings up, slowed them down or added background noise. Mixup is an exception to keeping labels fixed: it trains on blends of two examples and blends their labels too.

Most libraries build this in. Keras has layers such as RandomFlip and RandomRotation. PyTorch’s torchvision can transform bounding boxes, masks and keypoints along with the image, so the labels move with the picture. Mixing transforms such as MixUp and CutMix work on whole batches because they combine pairs of images. Picking the changes used to be done by hand. AutoAugment searches for a policy of image operations, how often to apply them and how strongly. The search aims for the best accuracy on a validation set. torchvision ships policies learned on ImageNet, CIFAR-10 and SVHN.

2 · Why it exists

Big models trained on too few examples tend to memorise them instead of learning the pattern.

Not enough examplesThe best fix is more labelled data, but collecting it is not always possible.
Memorising the training setA model with many parameters can overfit, scoring well on its training examples and poorly on new ones.
Harmless variations confuse itThe same object can appear shifted, mirrored or lit differently, and the model should give the same answer each time.
3 · How it works

Follow one labelled photo through training.

One labelled photo becomes many training copies that all keep the label cat. Test photos skip augmentation.
  1. 1 · takeTake an example from the training set only, such as a photo labelled cat.
  2. 2 · changeApply a random, realistic change that should not alter the answer, such as a flip, a crop, a small rotation or a brightness shift.
  3. 3 · keepThe changed copy keeps the original label, so the model gets an extra correct example.
  4. 4 · repeatNext time the photo comes round in training, a fresh random change is drawn.
  5. 5 · testValidation and test images are left unchanged, so scores still describe real inputs.

Only the training set is augmented. Evaluation and prediction use the images as they are.

4 · Where it's used
WhoWhat they askWhat it works with
Photo app team“Will our flower classifier still work on tilted or dim photos?”Training photos, randomly flipped, rotated and brightened
Speech recognition team“How do we stop the model memorising our limited recordings?”Spectrograms with blocks of time and frequency masked out
Support team with a small labelled set“Can we stretch a few hundred labelled tickets further?”Sentences with synonyms swapped in or words reordered
Object detection team“If we flip a street photo, do the box coordinates move too?”Images together with their bounding boxes and masks
5 · What it solves, and what it doesn't
solves
  • Reduces overfitting when data is scarce. AlexNet's authors said their network overfitted substantially without crops and flips.
  • Teaches invariances, so an object that is shifted, mirrored or relit still gets the same answer.
  • Costs little. Changes can be made on the fly during training, so no copies are stored on disk.
  • Works beyond photos. SpecAugment's masking reached state-of-the-art speech recognition results on LibriSpeech and Switchboard.
doesn't solve
  • It does not add truly new information. Copies of one image are highly interdependent.
  • A change can break the label. Horizontal flips help on CIFAR-10 photos but not on MNIST handwritten digits.
  • Heavy edits to text can change a sentence's meaning, leaving it with the wrong label.
  • The right changes depend on the dataset, and choosing them by hand takes expert knowledge and time.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsData augmentation, TensorFlow (Google) · read 27 Sept 2026
  2. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  3. docsTransforming images, videos, boxes and more (torchvision.transforms.v2), PyTorch · read 27 Sept 2026
  4. paperImageNet Classification with Deep Convolutional Neural Networks, Krizhevsky et al., NeurIPS 2012 · read 27 Sept 2026
  5. paperAutoAugment: Learning Augmentation Policies from Data, arXiv (Google Brain, Cubuk et al.; CVPR 2019) · read 27 Sept 2026
  6. papermixup: Beyond Empirical Risk Minimization, arXiv (Zhang et al.; ICLR 2018) · read 27 Sept 2026
  7. paperSpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition, arXiv (Park et al.; Interspeech 2019) · read 27 Sept 2026
  8. officialSpecAugment: A New Data Augmentation Method for Automatic Speech Recognition, Google Research · read 27 Sept 2026
  9. paperEDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks, arXiv (Wei and Zou; EMNLP-IJCNLP 2019) · read 27 Sept 2026