Probability
A probability is a number from 0 to 1 that says how likely something is. Many machine learning problems need one as the answer.
A probability is a number from 0 to 1 that says how likely something is. It can be read as how often something happens over many tries, or as how strongly you believe it. An event is any outcome you can check for, such as “this email is spam”. The probabilities of all the possible outcomes add up to exactly 1. So the chance that an event does not happen, called its complement, is 1 minus the chance that it does.
A probability distribution lists every possible outcome with its probability. A height can come out as endlessly many values, so the chance of hitting one exact figure, say 170.000 cm, is zero. Instead you ask how likely it is to land between two limits, and the answer is the size of the region under the curve across that stretch.
Joint probability is the chance that two events both happen. With counts, it is the number of cases where both happen divided by the total. Conditional probability, written P(A | B), is the chance of one event once you know another has happened. You keep only the cases where the second event is true, and ask what fraction of them have the first. Two events are independent when knowing one tells you nothing new about the other. Then their joint probability is the two probabilities multiplied. In the illustrative inbox, spam and “free” are not independent, since 0.30 × 0.25 = 0.075, while 20 of the 100 emails, or 0.20, are both.
Logistic regression passes a weighted sum of an email’s features through the sigmoid function, which always returns a value between 0 and 1. An output of 0.932 means a 93.2% chance that the email is spam. With more than two classes, the softmax function gives each class a probability, and these add up to exactly 1. A language model estimates the probability of the next token, a word or piece of a word, given the tokens before it. The chance of a whole text is then written as a chain of these conditional probabilities multiplied together.
A good probability should match how often the model is actually right. If a well calibrated model gives 100 emails a score near 0.8, about 80 of them should turn out to be spam. A 2017 study found that the deep networks of that time often failed that test, unlike the networks built ten years before them. The same paper’s repair was tiny: temperature scaling, which tunes just one number, brought the scores back in line on most of the data sets it tried.
Machine learning is uncertain about its data, its models and its answers.
Work out one conditional probability from a pile of counted emails.
- 1 · countSort 100 emails by spam or not and by whether they contain free, and count each pair.
- 2 · divideDivide a row or column total by all 100, such as P(spam) = 30 ÷ 100 = 0.30.
- 3 · restrictKeep only the 25 emails with free and ask what fraction are spam, which is 20 ÷ 25 = 0.80.
- 4 · decideA classifier reports such a probability for a new email, and a 0.5 threshold turns 0.80 into spam.
A condition shrinks the group you count: spam is 30 in 100 overall, but 20 in 25 among emails with free.
| Who | What they ask | What it works with |
|---|---|---|
| Email provider | “How likely is this message to be spam?” | One incoming email, scored as a probability of spam |
| Chat assistant team | “Which word should the model write next?” | A probability for every possible next token |
| Photo app team | “Is this picture a dog, a cat or a horse?” | Softmax probabilities over the possible labels |
| Hospital data team | “When our model says 0.8, is it right 80% of the time?” | Past predictions checked against real outcomes |
- Gives one scale, from 0 to 1, for how likely any event is.
- Lets a model say how sure it is, not just which label it picked.
- Lets people move a threshold to change decisions without retraining the model.
- Breaks the chance of a whole sentence into one conditional step per token.
- A model's probabilities are not automatically honest. A 2017 study found that many modern neural networks gave scores that did not match how often they were right.
- P(spam | free) is not P(free | spam). Swapping the two needs Bayes' theorem.
- Finding no straight-line link between two quantities does not prove they are independent.
- A probability alone makes no decision. Someone still has to choose the threshold.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMathematics for Machine Learning, chapter 6: Probability and Distributions, Deisenroth, Faisal and Ong, Cambridge University Press · read 27 Sept 2026
- officialWhat is a Probability Distribution (NIST/SEMATECH e-Handbook of Statistical Methods, 1.3.6.1), NIST/SEMATECH · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsLogistic regression: Calculating a probability with the sigmoid function (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsThresholds and the confusion matrix (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsProbability calibration, scikit-learn · read 27 Sept 2026
- paperOn Calibration of Modern Neural Networks, Guo, Pleiss, Sun and Weinberger, ICML 2017 · read 27 Sept 2026
- paperLanguage Models are Unsupervised Multitask Learners, OpenAI · read 27 Sept 2026