Skip to content
Launchpad Library logo
← All glossary terms

Philosophies & Movements

Mesa-Optimization (Inner Alignment)

What happens when training a model produces a system that is itself pursuing a goal — and that internal goal may differ from the one the trainers were optimizing for. “Inner alignment” is the problem of making the two match.

Origin

Coined in the 2019 paper “Risks from Learned Optimization in Advanced Machine Learning Systems” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant.

Read the source →