How It Works
Mechanistic Interpretability
Research that tries to reverse-engineer what is happening inside a neural network: which internal features and circuits of neurons carry out a behavior, rather than just watching what goes in and comes out. The goal is to read a model's reasoning directly, for example to spot deception or hidden goals. The field is still young and can explain only small pieces of today's large models.
Origin
The label is usually credited to Chris Olah, who used "mechanistic interpretability" for his circuits research at OpenAI and later Anthropic. The "Zoom In: An Introduction to Circuits" article (Olah and colleagues, Distill, 2020) set out the approach. The broader goal of interpreting neural networks is much older, so the credit is for the name and framing, not the idea.
