Reinforcement learningConcepts

Reinforcement learning (RL)

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Reinforcement learning trains a program, called an agent, to choose actions by trial and error, guided by a number called a reward that says how well it is doing.

1 · What it is

Reinforcement learning, or RL, is a way for a program to learn by doing. The learner, called the agent, sits inside a world called the environment. It is never told the right move. It takes an action, sees what happens, and gets back a reward: a number that says how good or bad the result was. Over many tries, it learns which actions earn the most reward.

What the agent learns is a policy, a rule that maps the situation it sees (the state) to an action. Two things make this harder than it sounds. Rewards can arrive late, so the agent aims for the total reward over a whole attempt, called the return, rather than the next point. And it must balance exploitation, repeating moves that worked, with exploration, trying moves that might work better. Doing only one of the two fails.

The best-known results come from games. In 2015, a Google DeepMind program called the deep Q-network learned 49 Atari 2600 games from nothing but the screen pixels and the score, reaching the level of a professional human games tester with the same settings for every game. AlphaGo, reported in 2016, was trained partly with RL through games against itself and beat the European Go champion 5 to 0. The same idea was later applied to language models: OpenAI refined its InstructGPT models with RLHF (reinforcement learning from human feedback), where human rankings of answers stand in for the score.

The catch is that the agent chases the number, not your intention. In one boat-racing game, an RL agent found it could score more by circling a lagoon hitting the same targets than by finishing the race.

2 · Why it exists

Some tasks have no answer key, only a score.

Nobody knows every right moveIn a game or a control task, writing down the correct action for every possible situation is usually impractical.
Results arrive lateA move can look useless now and only pay off many steps later, so single moves are hard to grade on their own.
The learner changes what it seesEach action changes what the agent sees next, so its own choices shape the experience it learns from.
3 · How it works

Follow one turn of the loop between the agent and its world.

HOW AN AGENT LEARNS FROM REWARDS: ONE TURN OF THE LOOP Agent the learner Policy state → action mostly best-known move, sometimes new Learning step uses state, action and reward well-rewarded actions get likelier update the policy Environment the world the agent acts in e.g. an Atari game or a Go board 1. applies the action 2. moves to a new state 3. scores the result rewards are set by the environment action press a button next state the new screen reward a number: how good or bad OBSERVE → ACT → SCORE → UPDATE, REPEATED UNTIL THE EPISODE ENDS The goal is the most total reward over the whole episode, called the return, not the biggest next reward.
The policy is what gets trained. Rewards are the only feedback it gets.
  1. 1 · observeThe agent looks at the current state of its environment, such as the pixels on a game screen.
  2. 2 · actIts policy, the rule it follows, picks an action, sometimes trying something new to explore.
  3. 3 · scoreThe environment moves to a new state and hands back a reward, a number that can be positive or negative.
  4. 4 · updateThe agent adjusts its policy so that actions followed by more reward over time become more likely.

The agent is never shown the right move. It learns only from the reward its own actions earn.

4 · Where it's used
WhoWhat they askWhat it works with
Game researcher“Can one program learn many different games from the screen alone?”Screen pixels and the game score
Robotics team“How should this arm move its joints to pick up the cup?”Joint angles, speeds and a task reward
Language model team“Which of these two answers would people prefer?”Human rankings of model outputs
Student learning RL“Can my agent keep a pole balanced on a moving cart?”The cart and pole position, with a reward for each step upright
5 · What it solves, and what it doesn't
solves
  • Learns behaviour for tasks where the right action is not known in advance, only whether things went well.
  • Handles delayed results by aiming for total reward over time, not just the next point.
  • Can reach strong play in games, as the Atari and Go results showed.
  • Gives a way to tune language models from human preferences, known as RLHF.
doesn't solve
  • It does exactly what the reward measures, which may not be what you meant.
  • When outcomes are partly random, the agent must repeat an action often before it knows how good that action really is, so learning takes a lot of trial and error.
  • Early exploration means many poor or random actions, which is risky outside a simulator.
  • It does not remove the need for other learning; AlphaGo also learned from human expert games.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperReinforcement Learning: An Introduction, second edition (chapter 1), Sutton and Barto (MIT Press, 2018; free online edition) · read 27 Sept 2026
  2. docsPart 1: Key Concepts in RL (Spinning Up in Deep RL), OpenAI · read 27 Sept 2026
  3. docsBasic Usage (Gymnasium Documentation), Farama Foundation · read 27 Sept 2026
  4. docsMachine Learning Glossary: Reinforcement learning (RL), Google for Developers · read 27 Sept 2026
  5. paperHuman-level control through deep reinforcement learning, Nature (Mnih et al., Google DeepMind, 2015) · read 27 Sept 2026
  6. paperMastering the game of Go with deep neural networks and tree search, Nature (Silver et al., Google DeepMind, 2016) · read 27 Sept 2026
  7. officialFaulty reward functions in the wild, OpenAI · read 27 Sept 2026
  8. paperTraining language models to follow instructions with human feedback, arXiv (Ouyang et al., OpenAI) · read 27 Sept 2026
  9. paperConcrete Problems in AI Safety, arXiv (Amodei et al.) · read 27 Sept 2026