Concepts

Q-learning

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Q-learning improves an estimate of each state-action value from sampled rewards and the best estimated value at the next state.

1 · What it is

Q-learning estimates action values for choices in each state. A sampled transition supplies the current state, chosen action, reward and next state.

The update target is the reward plus gamma times the largest estimated action value in the next state. The difference between that target and the current Q value determines the update.

The tabular convergence result assumes that every action is sampled repeatedly in every state and that action values are represented discretely. An epsilon-greedy policy usually selects a greedy action and sometimes selects an action at random. With neural function approximation, experience replay can smooth the training distribution. A maximum over uncertain estimates can also create positive overestimation bias. Double DQN separates action selection from evaluation in its target.

2 · Why it exists

An agent can learn action values from interaction without constructing an environment model.

Delayed rewardAn action value represents the maximum expected return after taking an action in a state.
Sampled experienceThe familiar Q-learning update replaces expectations with samples from behavior and environment distributions.
Bootstrapped targetThe update uses the largest estimated action value at the resulting state.
3 · How it works

Follow one transition through the Q-learning update.

One sampled transition updates the value of the action that was taken.
  1. 1 · observeRecord the state, selected action, reward and next state.
  2. 2 · targetAdd the reward to the discounted maximum action value at the next state.
  3. 3 · compareSubtract the current state-action value to obtain the update error.
  4. 4 · updateMove the selected Q value toward the target.
4 · Where it's used
WhoWhat they askWhat it works with
Control engineer“Which action has the highest estimated future return here?”Q values for the current state
RL researcher“Does a function approximator estimate the action-value function?”Q-network outputs
Training engineer“Is exploration sampling every relevant action?”Behavior-policy action counts
5 · What it solves, and what it doesn't
solves
  • The original result describes Q-learning as an incremental dynamic-programming method.
  • Deep Q-learning uses a neural network to estimate future rewards from raw pixels.
  • Experience replay samples previous transitions to smooth the training distribution.
doesn't solve
  • The convergence theorem requires repeated sampling of every action in every state and discrete action values.
  • A maximum over noisy value estimates can introduce positive overestimation bias.
  • Nonlinear function approximation and off-policy learning can make a Q-network diverge.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperQ-learning, Watkins and Dayan · read 28 Sept 2026
  2. paperPlaying Atari with Deep Reinforcement Learning, Mnih et al. · read 28 Sept 2026
  3. paperDouble Q-learning, van Hasselt · read 28 Sept 2026
  4. paperA Tutorial on Thompson Sampling, Russo et al. · read 28 Sept 2026
  5. paperDeep Reinforcement Learning with Double Q-learning, van Hasselt, Guez and Silver · read 28 Sept 2026