Q-learning
Q-learning improves an estimate of each state-action value from sampled rewards and the best estimated value at the next state.
Q-learning estimates action values for choices in each state. A sampled transition supplies the current state, chosen action, reward and next state.
The update target is the reward plus gamma times the largest estimated action value in the next state. The difference between that target and the current Q value determines the update.
The tabular convergence result assumes that every action is sampled repeatedly in every state and that action values are represented discretely. An epsilon-greedy policy usually selects a greedy action and sometimes selects an action at random. With neural function approximation, experience replay can smooth the training distribution. A maximum over uncertain estimates can also create positive overestimation bias. Double DQN separates action selection from evaluation in its target.
An agent can learn action values from interaction without constructing an environment model.
Follow one transition through the Q-learning update.
- 1 · observeRecord the state, selected action, reward and next state.
- 2 · targetAdd the reward to the discounted maximum action value at the next state.
- 3 · compareSubtract the current state-action value to obtain the update error.
- 4 · updateMove the selected Q value toward the target.
| Who | What they ask | What it works with |
|---|---|---|
| Control engineer | “Which action has the highest estimated future return here?” | Q values for the current state |
| RL researcher | “Does a function approximator estimate the action-value function?” | Q-network outputs |
| Training engineer | “Is exploration sampling every relevant action?” | Behavior-policy action counts |
- The original result describes Q-learning as an incremental dynamic-programming method.
- Deep Q-learning uses a neural network to estimate future rewards from raw pixels.
- Experience replay samples previous transitions to smooth the training distribution.
- The convergence theorem requires repeated sampling of every action in every state and discrete action values.
- A maximum over noisy value estimates can introduce positive overestimation bias.
- Nonlinear function approximation and off-policy learning can make a Q-network diverge.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperQ-learning, Watkins and Dayan · read 28 Sept 2026
- paperPlaying Atari with Deep Reinforcement Learning, Mnih et al. · read 28 Sept 2026
- paperDouble Q-learning, van Hasselt · read 28 Sept 2026
- paperA Tutorial on Thompson Sampling, Russo et al. · read 28 Sept 2026
- paperDeep Reinforcement Learning with Double Q-learning, van Hasselt, Guez and Silver · read 28 Sept 2026