Imbalanced data
Imbalanced data is a dataset where one class is far rarer than another, such as a few fraud cases among many ordinary card payments.
Imbalanced data, or a class-imbalanced dataset, is a set of labelled examples where one label is far more common than another. The common label is called the majority class and the rare one the minority class. This is normal in real data. Fraud can be under 0.1% of card transactions, and patients with a rare virus might be under 0.01% of a medical dataset. Datasets with more than two classes can be imbalanced too.
Training struggles because each small batch a model learns from needs examples of both classes. In one of Google’s examples, 2 rare examples sit among 202, so a batch of 20 usually holds none. What matters is how many minority examples there are, not the size of the whole dataset. Scoring struggles as well. A model that always predicts the majority class can post a very high accuracy with no predictive power. So evaluators tend to use precision and recall instead.
Fixes act at different points. Google’s course recommends downsampling the majority class and then upweighting it, shown in the figure. Upweighting means the loss, a measure of how far each prediction is from its label, counts each majority mistake more heavily. Skip it and the average prediction drifts from the average label. Class weights reach a similar place without dropping any rows. In scikit-learn, the balanced option weights each class in inverse proportion to how often it appears, so with 990 negatives and 10 positives each positive counts 99 times as much. Simple oversampling copies rare examples at random, and random undersampling deletes common ones. SMOTE, from a 2002 paper, instead makes synthetic rare examples on the line between one minority example and a nearby one. After training, the decision threshold can move too: scikit-learn predicts positive above a probability of 0.5 by default, but a tuned threshold can sit much lower, around 0.02 in one of its examples. To judge the result, a 2015 study found precision-recall plots more informative than ROC plots on imbalanced data. PR AUC, the area under that curve, sums it up: a random guesser scores the positive share, 0.01 at a 99 to 1 split.
When one class is rare, both training and scoring go wrong.
Follow one lopsided training set through downsampling and upweighting.
- 1 · countStart from the real data, here an illustrative 990 negatives and 10 positives, a 99 to 1 split.
- 2 · downsampleTrain on only a small share of the majority class, here 1 in 10, so rare examples turn up in far more batches.
- 3 · upweightMultiply the loss on each kept majority example by the same factor, 10, so the model still learns how common each class really is.
- 4 · trainWith rare examples showing up in batch after batch, the model settles on good weights sooner.
- 5 · tuneTreat the two factors as hyperparameters and try several values, since no single ratio is always best.
Downsampling fixes what the batches show. Upweighting restores how common each class really is.
| Who | What they ask | What it works with |
|---|---|---|
| Card fraud team | “Why does our model almost never flag a purchase?” | How many fraud cases land in each training batch |
| Hospital screening team | “Should we flag a scan at a 10% chance of disease instead of 50%?” | The decision threshold for the rare diagnosis |
| Factory quality team | “We have 40 photos of cracked parts and 50,000 good ones, so what now?” | Oversampled or synthetic examples of the defect class |
| Model evaluators | “Which model is better at finding the rare class?” | Precision-recall curves and PR AUC instead of accuracy |
- After downsampling, far fewer batches arrive with too few rare examples to learn from.
- Upweighting after downsampling keeps the model's sense of how common each class really is.
- Lowering the decision threshold catches more rare cases, at the cost of more false alarms.
- Precision-recall curves show rare-class performance that ROC plots can make look more reliable than it is.
- It cannot invent real evidence. A huge dataset can still fall short when the rare class has only a handful of examples.
- Downsampling alone leaves the model thinking rare cases are more common than they are, until upweighting corrects it.
- Copying rare examples to oversample them can make the model overfit to those few cases.
- There is no fixed best ratio. The downsampling and upweighting factors have to be found by experiment.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsDatasets: Class-imbalanced datasets (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsLogisticRegression (scikit-learn API reference), scikit-learn developers · read 27 Sept 2026
- docs3.3. Tuning the decision threshold for class prediction, scikit-learn developers · read 27 Sept 2026
- docs2. Over-sampling (imbalanced-learn user guide), imbalanced-learn developers · read 27 Sept 2026
- docs3. Under-sampling (imbalanced-learn user guide), imbalanced-learn developers · read 27 Sept 2026
- paperSMOTE: Synthetic Minority Over-sampling Technique, Chawla et al., JAIR 2002 · read 27 Sept 2026
- paperThe Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets, Saito and Rehmsmeier, PLOS ONE 2015 · read 27 Sept 2026