Concepts

Bayesian inference

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

Bayesian inference treats an unknown number as uncertain, starts from a belief about it, and uses Bayes' theorem to update that belief as data arrives.

1 · What it is

Bayesian inference treats an unknown quantity, such as the share of people who click an ad, as uncertain rather than fixed. You begin with a prior, a distribution that says how plausible each value seems before looking at the data. The likelihood scores how well each possible value explains the data you actually saw. Multiplying the two and rescaling so the total is 1 gives the posterior, your updated belief. That step is Bayes’ theorem, which the probability entry introduces.

The result is a whole distribution, not one number. You can summarise it with a mean and a credible interval, a range that holds a chosen share of the posterior probability, such as 90%. Because the unknown is treated as random, that interval is a direct probability statement about it. Classical, or frequentist, statistics instead treats the unknown as a fixed constant. Its confidence interval describes how often the method’s intervals would capture the true value, which is a different kind of statement.

Some pairs of prior and likelihood fit together so neatly that the posterior has the same form as the prior, a property called conjugacy. A Beta prior on a rate, updated with success and failure counts, is the classic case. The posterior is again a Beta, with the successes added to its first number and the failures to its second. Its mean always lies between the prior mean and the observed rate. As the counts grow, the prior matters less and less.

Once a model gets complicated, there may be no conjugate prior to use at all. Software then draws many random samples from the posterior and summarises those instead. A common family of such methods is Markov chain Monte Carlo, or MCMC. The Stan software, for example, runs an MCMC sampler called NUTS by default. In PyMC, another such tool, the model is ordinary Python, not a separate modelling language. Variational inference swaps sampling for optimisation. It fixes a set of candidate distributions, then searches that set for the one nearest the posterior. It usually finishes sooner than sampling. One review of the method says its behaviour is still only partly understood.

Bayesian ideas turn up across machine learning. Bayesian regression, for example, learns how strong its regularisation should be from the data instead of having it set by hand. The usual penalty in ridge regression matches a Bayesian estimate under a Gaussian prior. Bayesian optimisation tunes a model’s hyperparameters by treating how well the model performs as a draw from a Gaussian process. The posterior from past runs then guides which settings to try next. A 2012 study reported that it could match or beat human experts at tuning. The approach also suits A/B tests, where results arrive over time and a choice must be made.

Choosing the model, prior included, is a big hurdle in much Bayesian work. Part of checking the work is asking how much the answer shifts if those modelling choices change. Computation is the other cost, since fitting can be time consuming.

2 · Why it exists

A single best guess hides how sure you are, and classical methods leave out what you already knew.

One number hides doubtAn estimate on its own says nothing about how uncertain it is. Reporting that uncertainty is nearly always needed.
Past knowledge left outClassical analysis usually restricts itself to the current data. Prior knowledge mostly just helps pick which model to fit.
Data keeps arrivingWhen results come in batches, you want to fold in each new batch without redoing the whole calculation.
3 · How it works

Follow one ad's click rate as the views come in.

Illustrative numbers. With a Beta prior and click counts, the key step just adds clicks and misses to the prior's two numbers.
  1. 1 · priorPick a prior, a distribution saying how plausible each value is before any data; here Beta(1, 1) treats every click rate alike.
  2. 2 · likelihoodWrite the likelihood, which says how probable the observed data, 3 clicks in 10 views, would be under each possible rate.
  3. 3 · updateMultiply prior by likelihood to get the posterior; for a Beta prior this adds the clicks and misses, giving Beta(4, 8).
  4. 4 · summariseReport the whole posterior, such as its mean of about 0.33 and a credible interval around it.
  5. 5 · repeatUse that posterior as the prior for the next batch of views, and the belief narrows further.

In one line, posterior ∝ prior × likelihood: each possible value is reweighted by how well it explains the data.

4 · Where it's used
WhoWhat they askWhat it works with
Product team“Does version B of the page get more clicks than version A?”Click counts from an A/B test, with a prior on each version's rate
Machine learning engineer“Which learning rate should the next training run try?”Scores from earlier runs, modelled with a Gaussian process
Reliability engineer“How long does this system run between failures?”New test failures combined with earlier test results
Data scientist“How strong should the regularisation in this regression be?”Bayesian ridge regression, which estimates it from the data
5 · What it solves, and what it doesn't
solves
  • Gives a full distribution for each unknown, so the answer carries its own uncertainty.
  • A credible interval is a direct probability statement about the unknown.
  • New data can be added by treating the last posterior as the next prior.
  • With enough data, the prior's influence fades away.
doesn't solve
  • It does not say which prior is right, and different ways of setting one can give different results.
  • A prior built on inaccurate information can lead to misleading conclusions.
  • Fitting Bayesian models can be time consuming.
  • It does not check the model itself. You still have to test how well it fits the data.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperBayesian Data Analysis, third edition (free electronic edition), Gelman, Carlin, Stern, Dunson, Vehtari and Rubin, CRC Press · read 27 Sept 2026
  2. paperProbabilistic Machine Learning: An Introduction (online version), Murphy, MIT Press · read 27 Sept 2026
  3. paperProbabilistic Machine Learning: Advanced Topics (online version), 34.3 A/B testing, Murphy, MIT Press · read 27 Sept 2026
  4. officialHow can Bayesian methodology be used for reliability evaluation? (NIST/SEMATECH e-Handbook of Statistical Methods, 8.1.10), NIST/SEMATECH · read 27 Sept 2026
  5. docsMCMC Sampling (Stan Reference Manual), Stan Development Team · read 27 Sept 2026
  6. docsIntroductory Overview of PyMC, PyMC · read 27 Sept 2026
  7. paperVariational Inference: A Review for Statisticians, Blei, Kucukelbir and McAuliffe · read 27 Sept 2026
  8. paperPractical Bayesian Optimization of Machine Learning Algorithms, Snoek, Larochelle and Adams, NeurIPS 2012 · read 27 Sept 2026
  9. docs1.1. Linear Models: Bayesian Regression (scikit-learn user guide), scikit-learn · read 27 Sept 2026