Concepts

Mechanistic interpretability

Reference entrySafety, ethics and evaluation
Reference entryMechanistic interpretability

Mechanistic interpretability investigates neural networks by identifying internal components, representations, and computations that causally produce particular observed behaviors.

Where it sits