Skip to content
Launchpad Library logo
← All glossary terms

Philosophies & Movements

Corrigibility

The property of an AI system that does not resist being corrected, shut down, or modified by its operators — and does not try to manipulate them into leaving it alone.

Origin

Named and formalized in the 2015 paper “Corrigibility” by Nate Soares, Benja Fallenstein, Eliezer Yudkowsky (profiled on our Who’s Who page) and Stuart Armstrong, presented at the AAAI-15 AI and Ethics workshop; the paper credits Robert Miles with suggesting the term.

Read the source →