Skip to content
Launchpad Library logo
← All glossary terms

How It Works

Mechanistic Interpretability

Research that tries to reverse-engineer what is happening inside a neural network: which internal features and circuits of neurons carry out a behavior, rather than just watching what goes in and comes out. The goal is to read a model's reasoning directly, for example to spot deception or hidden goals. The field is still young and can explain only small pieces of today's large models.

Origin

The label is usually credited to Chris Olah, who used "mechanistic interpretability" for his circuits research at OpenAI and later Anthropic. The "Zoom In: An Introduction to Circuits" article (Olah and colleagues, Distill, 2020) set out the approach. The broader goal of interpreting neural networks is much older, so the credit is for the name and framing, not the idea.

Read the source →