Reinforcement learning (RL)
Reinforcement learning trains a program, called an agent, to choose actions by trial and error, guided by a number called a reward that says how well it is doing.
Reinforcement learning, or RL, is a way for a program to learn by doing. The learner, called the agent, sits inside a world called the environment. It is never told the right move. It takes an action, sees what happens, and gets back a reward: a number that says how good or bad the result was. Over many tries, it learns which actions earn the most reward.
What the agent learns is a policy, a rule that maps the situation it sees (the state) to an action. Two things make this harder than it sounds. Rewards can arrive late, so the agent aims for the total reward over a whole attempt, called the return, rather than the next point. And it must balance exploitation, repeating moves that worked, with exploration, trying moves that might work better. Doing only one of the two fails.
The best-known results come from games. In 2015, a Google DeepMind program called the deep Q-network learned 49 Atari 2600 games from nothing but the screen pixels and the score, reaching the level of a professional human games tester with the same settings for every game. AlphaGo, reported in 2016, was trained partly with RL through games against itself and beat the European Go champion 5 to 0. The same idea was later applied to language models: OpenAI refined its InstructGPT models with RLHF (reinforcement learning from human feedback), where human rankings of answers stand in for the score.
The catch is that the agent chases the number, not your intention. In one boat-racing game, an RL agent found it could score more by circling a lagoon hitting the same targets than by finishing the race.
Some tasks have no answer key, only a score.
Follow one turn of the loop between the agent and its world.
- 1 · observeThe agent looks at the current state of its environment, such as the pixels on a game screen.
- 2 · actIts policy, the rule it follows, picks an action, sometimes trying something new to explore.
- 3 · scoreThe environment moves to a new state and hands back a reward, a number that can be positive or negative.
- 4 · updateThe agent adjusts its policy so that actions followed by more reward over time become more likely.
The agent is never shown the right move. It learns only from the reward its own actions earn.
| Who | What they ask | What it works with |
|---|---|---|
| Game researcher | “Can one program learn many different games from the screen alone?” | Screen pixels and the game score |
| Robotics team | “How should this arm move its joints to pick up the cup?” | Joint angles, speeds and a task reward |
| Language model team | “Which of these two answers would people prefer?” | Human rankings of model outputs |
| Student learning RL | “Can my agent keep a pole balanced on a moving cart?” | The cart and pole position, with a reward for each step upright |
- Learns behaviour for tasks where the right action is not known in advance, only whether things went well.
- Handles delayed results by aiming for total reward over time, not just the next point.
- Can reach strong play in games, as the Atari and Go results showed.
- Gives a way to tune language models from human preferences, known as RLHF.
- It does exactly what the reward measures, which may not be what you meant.
- When outcomes are partly random, the agent must repeat an action often before it knows how good that action really is, so learning takes a lot of trial and error.
- Early exploration means many poor or random actions, which is risky outside a simulator.
- It does not remove the need for other learning; AlphaGo also learned from human expert games.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperReinforcement Learning: An Introduction, second edition (chapter 1), Sutton and Barto (MIT Press, 2018; free online edition) · read 27 Sept 2026
- docsPart 1: Key Concepts in RL (Spinning Up in Deep RL), OpenAI · read 27 Sept 2026
- docsBasic Usage (Gymnasium Documentation), Farama Foundation · read 27 Sept 2026
- docsMachine Learning Glossary: Reinforcement learning (RL), Google for Developers · read 27 Sept 2026
- paperHuman-level control through deep reinforcement learning, Nature (Mnih et al., Google DeepMind, 2015) · read 27 Sept 2026
- paperMastering the game of Go with deep neural networks and tree search, Nature (Silver et al., Google DeepMind, 2016) · read 27 Sept 2026
- officialFaulty reward functions in the wild, OpenAI · read 27 Sept 2026
- paperTraining language models to follow instructions with human feedback, arXiv (Ouyang et al., OpenAI) · read 27 Sept 2026
- paperConcrete Problems in AI Safety, arXiv (Amodei et al.) · read 27 Sept 2026