Continual learning
Continual learning trains a model on a stream of changing data while measuring whether later learning damages or helps earlier tasks.
The model observes examples from a sequence of tasks. Metrics can evaluate both test accuracy and transfer across those tasks.
EWC selectively slows learning on weights important for earlier tasks. GEM alleviates forgetting while allowing transfer to previous tasks. Learning without Forgetting uses new-task data while preserving original capabilities. iCaRL adds new classes progressively while keeping data for only a small number of classes at once.
Continual learning therefore evaluates two outcomes together: learning from the current stream and retaining performance acquired before it.
A model that learns tasks one after another can lose performance on knowledge it acquired earlier.
Follow one model through three tasks.
- 1 · observeThe model receives examples from a sequence of tasks or a changing data stream.
- 2 · updateTraining uses the current examples and may also use a retention mechanism.
- 3 · retainOne method selectively slows learning on weights that are important for earlier tasks.
- 4 · evaluateTests after each task measure current accuracy and transfer to earlier tasks.
The defining test is not only whether the model learns next, but what happens to performance from before.
| Who | What they ask | What it works with |
|---|---|---|
| Vision team | “Can new classes be added without retraining on the complete archive?” | Class-incremental accuracy after each addition |
| Robotics lab | “Can an agent keep skills while its environment changes?” | Performance across a stream of experiences |
| Research engineer | “Which strategy reduces forgetting on this benchmark?” | Accuracy and transfer metrics across tasks |
- It provides a setting for learning from non-stationary streams.
- It makes retention and transfer measurable across a sequence of tasks.
- iCaRL is a training strategy that adds new classes progressively.
- No single method guarantees that all earlier performance will be preserved.
- Keeping examples, extra model state or task-specific constraints adds storage or computation.
- Results remain sensitive to the task sequence, evaluation protocol and available past data.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperOvercoming catastrophic forgetting in neural networks, Kirkpatrick et al. · read 28 Sept 2026
- paperGradient Episodic Memory for Continual Learning, Lopez-Paz and Ranzato · read 28 Sept 2026
- paperLearning without Forgetting, Li and Hoiem · read 28 Sept 2026
- paperiCaRL: Incremental Classifier and Representation Learning, Rebuffi et al. · read 28 Sept 2026
- paperAvalanche: an End-to-End Library for Continual Learning, Lomonaco et al. · read 28 Sept 2026