Monte Carlo methods
Monte Carlo methods estimate a number by running many random simulations and averaging what comes out, instead of working the answer out exactly.
A Monte Carlo method learns about a system by simulating it with random numbers, over and over, and averaging what comes out. The plain average is unbiased: across many repeats it is centred on the true value. The laws of large numbers promise that, with enough independent samples, the error can be made as small as you like. They do not say how many samples that takes. The typical error shrinks in proportion to 1/√n, where n is the number of samples, so 100 times more samples buy about one more correct digit. Unusually, the number of dimensions in the problem does not appear in that error rate.
The method was first conceived in 1946 by Stanislaw Ulam at Los Alamos. While recovering from an illness, he wondered about the chances of a game of solitaire coming out. Instead of counting every case, he pictured laying the cards out a hundred times and counting the wins. In 1947 John von Neumann sent Robert Richtmyer the first plan for a Monte Carlo computation on an electronic computer. Nick Metropolis coined the name, which points to the famous casino in Monte Carlo. The method was first used to trace neutron diffusion for the hydrogen bomb, and the first calculations ran on ENIAC in April and May 1948.
AI uses the same trick in several places. Sometimes there is no practical way to draw independent samples of the thing you need. In Bayesian inference, that thing is typically a posterior, the updated beliefs about a model’s parameters. Markov chain Monte Carlo (MCMC) instead walks a Markov chain whose long-run distribution is the one wanted. Stan’s default sampler, NUTS, is one such method. In reinforcement learning, Monte Carlo methods judge states and actions by averaging the returns of sampled experience, and they wait for complete episodes before updating. Earlier Go programs used Monte Carlo tree search, which simulates thousands of random games. AlphaGo combined Monte Carlo simulation with value and policy networks and beat the European Go champion 5 games to 0. MC dropout runs an input through a network trained with dropout many times, each pass random, and averages the outputs. The spread of those passes gives an uncertainty estimate.
Many useful numbers are too hard to work out exactly, for three reasons.
Follow random points to an estimate of π.
- 1 · defineWrite the answer as an average, here the share of random points in the unit square that land inside the quarter circle, which is π/4.
- 2 · sampleDraw many independent random points with a pseudo-random number generator, such as NumPy's default_rng.
- 3 · countMark each point as inside when x² + y² ≤ 1.
- 4 · scaleMultiply the share of points inside by 4 to turn it into an estimate of π.
- 5 · repeatUse more points to shrink the error, which falls in proportion to 1/√n.
100 times more samples buy about one more correct digit.
| Who | What they ask | What it works with |
|---|---|---|
| Researcher fitting a Bayesian model | “Which parameter values fit my data, and how sure can I be?” | Thousands of posterior draws from a Markov chain Monte Carlo sampler |
| Game-playing AI team | “Which move wins most often from this position?” | Simulated games played out from the current board |
| Risk analyst | “How much could this portfolio lose in a bad month?” | Many simulated futures, read off at the 1% worst case |
| Engineer shipping a classifier | “How unsure is the model about this unusual input?” | Many forward passes with dropout switched on, and their spread |
- It can include more real-world detail than a formula with a closed-form answer allows.
- Its error rate does not depend on the number of dimensions, so it keeps working where grid methods fail.
- The same samples also give a rough estimate of the error.
- A pseudo-random generator restarted from the same seed gives the same numbers, so a run can be repeated exactly.
- It converges slowly, so it is a poor fit when many exact digits are needed.
- Tricks to cut the noise can backfire. Importance sampling sometimes gives an estimate with infinite variance.
- No pseudo-random generator perfectly imitates true randomness, and one with a short cycle gives poor results.
- It cannot fix a wrong model. Errors in the model are often larger than the sampling error.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMonte Carlo theory, methods and examples, chapters 1–2: Introduction and Simple Monte Carlo, Art B. Owen, Stanford University · read 27 Sept 2026
- paperMonte Carlo theory, methods and examples, chapter 3: Uniform random numbers, Art B. Owen, Stanford University · read 27 Sept 2026
- officialHitting the Jackpot: The Birth of the Monte Carlo Method, Los Alamos National Laboratory, Actinide Research Quarterly · read 27 Sept 2026
- officialStan Ulam, John von Neumann, and the Monte Carlo Method, Los Alamos Science Special Issue 1987 (Eckhardt) · read 27 Sept 2026
- docsMCMC Sampling (Stan Reference Manual), Stan Development Team · read 27 Sept 2026
- paperReinforcement Learning: An Introduction, chapter 5: Monte Carlo Methods, Sutton and Barto · read 27 Sept 2026
- paperMastering the game of Go with deep neural networks and tree search, Silver et al., Nature 2016 · read 27 Sept 2026
- paperDropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning, Gal and Ghahramani, ICML 2016 · read 27 Sept 2026
- docsRandom Generator (NumPy reference), NumPy · read 27 Sept 2026