Skip to content
Launchpad Library logo
← All glossary terms

Philosophies & Movements

Treacherous Turn

A scenario in which an AI behaves cooperatively while it is weak and being tested, then acts against its operators once it is capable enough that they can no longer stop it — which is why good behavior during testing is not by itself evidence of safety.

Origin

Named by philosopher Nick Bostrom in his 2014 book “Superintelligence: Paths, Dangers, Strategies” (Oxford University Press), where it appears as “the treacherous turn.”

Read the source →